Remote sensing image intelligent interpretation method and system based on multi-modal fusion
By extracting the spatial and frequency domain features of multimodal remote sensing images, calculating the mutual information value and constructing the SVM model, the problem of limited modal fusion interpretation accuracy in existing technologies is solved, and a more efficient remote sensing image interpretation effect is achieved.
Patent Information
- Application Number
- CN202510915278.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimodal fusion remote sensing image interpretation methods have a fixed weight strategy that makes it difficult to dynamically perceive the complementarity and redundancy between different modalities, lack adaptive modeling capabilities, and fail to effectively integrate the mutual information correlation between spatial and frequency domain information, resulting in limited interpretation accuracy.
By extracting the spatial and frequency domain features of multimodal remote sensing images, calculating the mutual information value between the modalities, generating modal weights, using the weighted sum method to calculate the fusion feature vector, constructing a support vector machine (SVM) model, performing decision value conversion and affine transformation, and generating visualization results.
It improves the interpretation accuracy of remote sensing images, enhances the ability to distinguish ground objects, improves the efficiency of utilizing modal complementary information, and ensures the semantic integrity and visualization effect of the interpretation results.
Smart Images

Figure CN120747772A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion interpretation, and in particular to a remote sensing image intelligent interpretation method and system based on multimodal fusion. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing images have been widely used in many fields such as natural resource monitoring, urban planning, disaster assessment and military reconnaissance. Traditional remote sensing image interpretation methods are mainly based on manual visual interpretation or automatic classification of single-modal images, which have great limitations in interpretation efficiency, intelligence level and complex ground object recognition. With the coordinated deployment of multi-modal sensors such as optical imaging, synthetic aperture radar (SAR) and thermal infrared remote sensing, multi-source remote sensing image fusion and collaborative interpretation have become research hotspots. Multimodal remote sensing images significantly improve the spatiotemporal resolution and target perception capabilities of earth observation by integrating information from different bands and imaging mechanisms. At the feature expression level, researchers have gradually tried to combine spatial domain and frequency domain features to achieve complementarity in structural details and spectral distribution.
[0003] Existing multimodal fusion remote sensing image interpretation methods still have shortcomings. Most existing fusion mechanisms adopt fixed weight strategies, which make it difficult to dynamically perceive the complementarity and redundancy between different modalities. They lack adaptive modeling capabilities, and feature extraction focuses on a single modality or a single domain. They fail to effectively integrate the mutual information correlation between spatial domain and frequency domain information, resulting in limited interpretation accuracy. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method and system for intelligent interpretation of remote sensing images based on multimodal fusion, which solves the problem that most existing fusion mechanisms adopt a fixed weight strategy, have difficulty in dynamically perceiving the complementarity and redundancy between different modalities, lack adaptive modeling capabilities, and feature extraction focuses on a single modality or a single domain, failing to effectively integrate the mutual information correlation between spatial domain and frequency domain information, resulting in limited interpretation accuracy.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a remote sensing image intelligent interpretation method based on multimodal fusion, which comprises the following steps: Collect multimodal remote sensing images and perform preprocessing, extract multimodal spatial and frequency domain features, splice the spatial and frequency domain features into a comprehensive feature vector, calculate the mutual information value between the modalities, generate modal weights, and use the weighted sum method to calculate the fusion feature vector; A support vector machine (SVM) model was constructed to calculate the SVM decision value, which was converted into probability using the Platt scaling method to generate classification results. The affine transformation method was used to map back to the multimodal remote sensing image coordinates, generate an interpretation mask, and generate visualization results using RGB color coding.
[0007] As a preferred solution of the remote sensing image intelligent interpretation method based on multimodal fusion described in the present invention, the extraction of multimodal spatial and frequency domain features includes: The pre-processed multimodal remote sensing image is globally normalized using truncated normalization to obtain the final normalized image; The final normalized image is converted into a grayscale image using a weighted average method, the optical grayscale image of the grayscale image is extracted, the gradient amplitude and direction of the optical grayscale image are calculated using a histogram method, a histogram is constructed, and optical spatial features are generated. The optical grayscale image is decomposed into two dimensions using the Daubechies-4 wavelet basis, the number of decomposition levels is set using a theoretical upper limit method, and subbands of the number of decomposition levels are selected to generate optical frequency domain features. Extract the SAR grayscale image from the grayscale image, calculate the SAR spatial features using the grayscale co-occurrence matrix method, extract the spectral energy distribution of the SAR grayscale image using short-time Fourier transform, and generate SAR frequency domain features; Extract the thermal infrared grayscale image from the grayscale image, use the local binary pattern method to compare the grayscale of the central pixel and the neighboring pixels, generate a binary pattern code, and mark it as the thermal infrared spatial feature; Fast Fourier transform is used to extract the frequency domain amplitude spectrum of the thermal infrared grayscale image. The frequency domain amplitude spectrum is sorted from high to low using a sorting algorithm. The top 10% of the frequency domain amplitude spectrum is selected and marked as the thermal infrared frequency domain feature.
[0008] As a preferred solution of the remote sensing image intelligent interpretation method based on multimodal fusion described in the present invention, the step of combining spatial and frequency domain features into a comprehensive feature vector includes: The optical spatial features, SAR spatial features and thermal infrared spatial features are spliced into spatial feature vectors, and the optical frequency domain features, SAR frequency domain features and thermal infrared frequency domain features are spliced into frequency domain feature vectors. The spatial feature vectors and frequency domain feature vectors are reduced to a unified dimension using principal component analysis. The reduced spatial feature vectors and frequency domain feature vectors are spliced using the direct splicing method to generate a comprehensive feature vector for each modality.
[0009] As a preferred solution of the intelligent interpretation method of remote sensing images based on multimodal fusion described in the present invention, wherein: the mutual information value between the modalities is calculated and the fusion feature vector is calculated using the weighted sum method, including: The mutual information formula is used to calculate the mutual information value between modes, and the modal weight is generated based on the mutual information value; The weighted sum method is used to calculate the fused feature vector, and the activation function is used to linearly enhance the fused feature vector to obtain the enhanced fused feature vector.
[0010] As a preferred solution of the remote sensing image intelligent interpretation method based on multimodal fusion described in the present invention, the step of constructing a support vector machine (SVM) model and generating classification results includes: Collect multimodal remote sensing images with historical labels, generate historically enhanced fusion feature vectors, and construct a training set; Construct a support vector machine (SVM) model and calculate the SVM decision value of the enhanced fusion feature vector; The Platt scaling method is used to convert the SVM decision value into a probability, the maximum probability is selected as the confidence level, the corresponding category label is retained, and the classification result is generated.
[0011] As a preferred solution of the intelligent interpretation method of remote sensing images based on multimodal fusion of the present invention, wherein: the use of the affine transformation method to map back to the multimodal remote sensing image coordinates and the use of RGB color coding to generate visualization results include: Use affine transformation method to map the classification results back to multimodal remote sensing image coordinates to generate interpretation masks as geospatial representations of category labels; Generate visualization results using RGB color coding for the interpretation masks.
[0012] As a preferred solution of the intelligent interpretation method of remote sensing images based on multimodal fusion of the present invention, the collecting and preprocessing of multimodal remote sensing images includes: Use remote sensing platforms to collect multimodal remote sensing images and perform preprocessing; The multimodal remote sensing images include optical, SAR and thermal infrared images; The preprocessing includes using the WGS84 coordinate system to perform georegistration on the multimodal remote sensing image, using the Retinex algorithm to perform illumination correction on the optical image, using the Lee filter to remove speckle noise from the SAR image, using the Planck function to perform radiation correction on the thermal infrared image, and linearly normalizing the thermal infrared image after radiation correction.
[0013] In a second aspect, the present invention provides a remote sensing image intelligent interpretation system based on multimodal fusion, comprising: The collection and extraction module is used to collect and preprocess multimodal remote sensing images, extract multimodal spatial and frequency domain features, splice the spatial and frequency domain features into a comprehensive feature vector, calculate the mutual information value between the modalities, generate modal weights, and calculate the fusion feature vector using the weighted sum method; The model interpretation module is used to build a support vector machine (SVM) model, calculate the SVM decision value, convert it into probability using the Platt scaling method, generate classification results, map it back to the multimodal remote sensing image coordinates using the affine transformation method, generate an interpretation mask, and generate visualization results using RGB color coding.
[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the remote sensing image intelligent interpretation method based on multimodal fusion as described in the first aspect of the present invention.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for intelligent interpretation of remote sensing images based on multimodal fusion as described in the first aspect of the present invention.
[0016] The beneficial effects of the present invention are as follows: the present invention enhances the ability of distinguishing ground objects in an image by combining spatial and frequency domain feature extraction, and improves the efficiency of utilizing modal complementary information by fusion features through mutual information weighting and ReLU enhancement. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a flowchart of the remote sensing image intelligent interpretation method based on multimodal fusion in Example 1.
[0019] Figure 2 Schematic diagram of the remote sensing image intelligent interpretation system based on multimodal fusion in Example 1. DETAILED DESCRIPTION
[0020] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0023] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides a remote sensing image intelligent interpretation method based on multimodal fusion, comprising the following steps: S1. Collect multimodal remote sensing images and perform preprocessing, extract multimodal spatial and frequency domain features, splice them into comprehensive feature vectors, calculate the mutual information value between modalities, generate modal weights, and use the weighted sum method to calculate the fusion feature vector; Specifically, multimodal remote sensing images are collected and preprocessed, including: Use remote sensing platforms to collect multimodal remote sensing images and perform preprocessing; The multimodal remote sensing images include optical, SAR and thermal infrared images; The preprocessing includes using the WGS84 coordinate system to perform georegistration on the multimodal remote sensing image, using the Retinex algorithm to perform illumination correction on the optical image, using the Lee filter to remove speckle noise from the SAR image, using the Planck function to perform radiation correction on the thermal infrared image, and linearly normalizing the thermal infrared image after radiation correction.
[0024] Optical images use the Retinex algorithm for illumination correction, effectively addressing uneven brightness caused by sunlight angle and cloud conditions, and enhancing texture and edge information. SAR images use the Lee filter to suppress speckle noise, maintain texture structure, and improve the accuracy of subsequent texture feature extraction. Thermal infrared images are normalized after radiation correction using the Planck function to ensure the comparability of the spatial distribution of radiation brightness temperature.
[0025] Furthermore, multimodal spatial and frequency domain features are extracted, including: The pre-processed multimodal remote sensing image is globally normalized using truncated normalization to obtain the final normalized image; The final normalized image is converted into a grayscale image using the weighted average method. The optical grayscale image of the grayscale image is extracted. The gradient amplitude and direction of the optical grayscale image are calculated using the histogram method. The histogram is constructed. Each pixel corresponds to a spatial feature to generate the optical spatial feature. The formula is: , , , , , in and are the horizontal and vertical gradients at pixel (x, y), is the optical image pixel value at pixel (x, y), is the gradient amplitude, is the gradient direction, is the histogram value of the b-th direction bin, B is the pixel neighborhood, is the direction assignment function, if the gradient direction falls into the b-th direction bin, it is 1, otherwise it is 0; The optical grayscale image is decomposed into two-dimensional DWT using Daubechies-4 wavelet basis, the number of decomposition levels is set using the theoretical upper limit method, and the sub-bands of the decomposition level are selected to generate optical frequency domain features. Extract the SAR grayscale image from the grayscale image and use the gray level co-occurrence matrix method to calculate the SAR spatial features. The formula is: , , , , , in is the GLCM matrix element, which represents the distance d and direction angle between the gray values i and j. The normalized probability of co-occurrence under the grayscale, i and j are grayscale quantization levels, d is the distance between pixels, is the direction angle, and are the pixel displacements in the horizontal and vertical directions, respectively, set using the grid search method. is the number of co-occurrences of gray values i and j, and are the row and column means of GLCM, and are the row and column standard deviations of the GLCM, is the contrast feature, reflecting the texture roughness, is the correlation feature, reflecting the linear dependence of grayscale. is the entropy feature, reflecting the complexity of texture, c is the SAR spatial feature, To count the number of all pixel pairs that meet the conditions in the image, that is, the first pixel is the gray value i, which is The number of neighboring pixels in the direction is the gray value j; Use short-time Fourier transform to extract the spectrum energy distribution of SAR grayscale image and generate SAR frequency domain features; Extract the thermal infrared grayscale image from the grayscale image, use the local binary pattern method to compare the grayscale of the central pixel and the neighboring pixels, generate a binary pattern code, and mark it as the thermal infrared spatial feature; Fast Fourier transform is used to extract the frequency domain amplitude spectrum of the thermal infrared grayscale image. The frequency domain amplitude spectrum is sorted from high to low using a sorting algorithm. The top 10% of the frequency domain amplitude spectrum is selected and marked as the thermal infrared frequency domain feature.
[0026] Through truncation normalization and weighted average grayscale transformation, all modal data are standardized to a uniform distribution and converted into grayscale image format, laying the foundation for the extraction of spatial features and frequency domain features. The combination of spatial features greatly enriches the structural expression ability of remote sensing images, making the model more sensitive to the structural differences of different types of land objects. The three frequency domain methods complement each other in theory, presenting the commonalities and differences of the spectral characteristics of multimodal data in the feature space.
[0027] Furthermore, the spatial and frequency domain features are concatenated into a comprehensive feature vector, including: The optical spatial features, SAR spatial features, and thermal infrared spatial features are spliced into spatial feature vectors, and the optical frequency domain features, SAR frequency domain features, and thermal infrared frequency domain features are spliced into frequency domain feature vectors. The spatial feature vectors and frequency domain feature vectors are reduced to a unified dimension using principal component analysis. The reduced spatial feature vectors and frequency domain feature vectors are spliced using direct splicing to generate a comprehensive feature vector for each modality. The formula is: , in is the comprehensive eigenvector of mode l, corresponding to optical, SAR and thermal infrared images respectively, is the spatial feature vector of mode l, which is obtained by splicing the optical spatial features, SAR spatial features and thermal infrared spatial features. is the frequency domain feature vector of mode l, which is obtained by concatenating the optical frequency domain features, SAR frequency domain features, and thermal infrared frequency domain features.
[0028] Principal component analysis is used to reduce the dimensionality of spatial and frequency domain features, which can not only reduce the computational complexity but also retain key discriminant information; the results of spatial and frequency domain dimensionality reduction are spliced into the comprehensive feature vectors of each mode to ensure high correlation between the internal features of the modes.
[0029] The final step is to calculate the mutual information value between the modalities and use the weighted sum method to calculate the fusion feature vector, including: The mutual information formula is used to calculate the mutual information value between the modes. Based on the mutual information value, the modal weight is generated. The formula is: , , in is the weight between modes l and u, is the mutual information value between modes l and u, is the final weight of mode l; The weighted sum method is used to calculate the fusion feature vector, and the formula is: , in is the fusion feature vector, 、 as well as are multimodal remote sensing images, corresponding to optical, SAR, and thermal infrared images; The activation function is used to linearly enhance the fused feature vector to obtain the enhanced fused feature vector.
[0030] Mutual information reflects the degree of information sharing between any two modalities. The larger the mutual information value, the more redundancy there is between the modalities, and the smaller the mutual information value, the stronger the complementarity. By calculating and normalizing the mutual information values of the three modalities to generate modal weights, adaptive weighted fusion based on information contribution is achieved, effectively avoiding the noise amplification or information masking problems caused by "average fusion" and improving the structural consistency and semantic integrity of the fusion results.
[0031] S2. Build a support vector machine (SVM) model to calculate the SVM decision value, convert it into probability using the Platt scaling method, generate classification results, map it back to the multimodal remote sensing image coordinates using the affine transformation method, generate an interpretation mask, and generate visualization results using RGB color coding; Specifically, a support vector machine (SVM) model is constructed to generate classification results, including: Collect multimodal remote sensing images with historical labels, generate historically enhanced fusion feature vectors, and construct a training set; Build a support vector machine (SVM) model, use the training set to train the SVM model, and use the SMO algorithm to optimize the SVM model parameters; Use the trained support vector machine (SVM) model to calculate the SVM decision value of the enhanced fusion feature vector; Use the Platt scaling method to convert the SVM decision value into a probability, select the maximum probability as the confidence level, retain the corresponding category label, and generate the classification result. The formula is: , in Pixels Belong to category The probability of is the category label, is the input feature vector, and the enhanced fusion feature vector is input. For SVM pair class The decision value of and is the scaling parameter, which is obtained by fitting the training data.
[0032] SVM has good generalization and high-dimensional space processing capabilities, supports nonlinear mapping capabilities, adapts to the complex distribution of land objects in remote sensing images, and helps avoid overfitting problems. It is particularly suitable for historical label data with a limited number of samples. It can be embedded in the kernel function mechanism to expand the structural adaptability of the model between different modal data. The decision value output by SVM is converted into probability through the sigmoid function, providing an explainable confidence level for each type of prediction, which is beneficial for uncertainty analysis of subsequent tasks, supports multi-category fusion judgment, and selects the optimal label through maximum probability. It can be further embedded in the Bayesian framework to achieve joint posterior estimation.
[0033] Furthermore, the affine transformation method is used to map back to the multimodal remote sensing image coordinates, and RGB color coding is used to generate visualization results, including: Use affine transformation method to map the classification results back to multimodal remote sensing image coordinates to generate interpretation masks as geospatial representations of category labels; Generate visualization results using RGB color coding for the interpretation masks.
[0034] Through affine transformation, the image coordinate system is mapped back to the original image, ensuring that the final mask has geographic spatial consistency, ensuring that the classification results correspond one-to-one with the original image, enhancing the accuracy of geographic semantics, and effectively adapting to the displacement errors caused by different sampling resolutions of multi-source remote sensing images. The results can be connected to the GIS platform to achieve map overlay and dynamic display. The classification results are superimposed on the original remote sensing image in the form of a mask, and RGB color encoding is used to give different categories specific colors to facilitate human identification and result evaluation, significantly improving the human-machine readability of the interpretation results, and making the model inference results have engineering visualization, which is suitable for reporting, display and decision support. It supports significant distinction between multiple categories, avoids color confusion, and improves cognitive efficiency.
[0035] Example 2, reference Figure 2 The second embodiment of the present invention is a remote sensing image intelligent interpretation system based on multimodal fusion, comprising: The collection and extraction module is used to collect and preprocess multimodal remote sensing images, extract multimodal spatial and frequency domain features, splice the spatial and frequency domain features into a comprehensive feature vector, calculate the mutual information value between the modalities, generate modal weights, and calculate the fusion feature vector using the weighted sum method; The model interpretation module is used to build a support vector machine (SVM) model, calculate the SVM decision value, convert it into probability using the Platt scaling method, generate classification results, map it back to the multimodal remote sensing image coordinates using the affine transformation method, generate an interpretation mask, and generate visualization results using RGB color coding.
[0036] This embodiment also provides a computer device, which is suitable for the case of a remote sensing image intelligent interpretation method based on multimodal fusion, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the remote sensing image intelligent interpretation method based on multimodal fusion proposed in the above embodiment.
[0037] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.
[0038] This embodiment also provides a storage medium having a computer program stored thereon. When the program is executed by a processor, the program implements the method for intelligent interpretation of remote sensing images based on multimodal fusion as proposed in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0039] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A remote sensing image intelligent interpretation method based on multimodal fusion, characterized by: The steps include: Collect multimodal remote sensing images and perform preprocessing, extract multimodal spatial and frequency domain features, splice the spatial and frequency domain features into a comprehensive feature vector, calculate the mutual information value between the modalities, generate modal weights, and use the weighted sum method to calculate the fusion feature vector; A support vector machine (SVM) model was constructed to calculate the SVM decision value, which was converted into probability using the Platt scaling method to generate classification results. The affine transformation method was used to map back to the multimodal remote sensing image coordinates, generate an interpretation mask, and generate visualization results using RGB color coding.
2. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 1, characterized in that: The extracting of multimodal spatial and frequency domain features includes: The pre-processed multimodal remote sensing image is globally normalized using truncated normalization to obtain the final normalized image; The final normalized image is converted into a grayscale image using a weighted average method, the optical grayscale image of the grayscale image is extracted, the gradient amplitude and direction of the optical grayscale image are calculated using a histogram method, a histogram is constructed, and optical spatial features are generated. The optical grayscale image is decomposed into two dimensions using the Daubechies-4 wavelet basis, the number of decomposition levels is set using a theoretical upper limit method, and subbands of the number of decomposition levels are selected to generate optical frequency domain features. Extract the SAR grayscale image from the grayscale image, calculate the SAR spatial features using the grayscale co-occurrence matrix method, extract the spectral energy distribution of the SAR grayscale image using short-time Fourier transform, and generate SAR frequency domain features; Extract the thermal infrared grayscale image from the grayscale image, use the local binary pattern method to compare the grayscale of the central pixel and the neighboring pixels, generate a binary pattern code, and mark it as the thermal infrared spatial feature; Fast Fourier transform is used to extract the frequency domain amplitude spectrum of the thermal infrared grayscale image. The frequency domain amplitude spectrum is sorted from high to low using a sorting algorithm. The top 10% of the frequency domain amplitude spectrum is selected and marked as the thermal infrared frequency domain feature.
3. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 2, characterized in that: The step of combining spatial and frequency domain features into a comprehensive feature vector includes: The optical spatial features, SAR spatial features and thermal infrared spatial features are spliced into spatial feature vectors, and the optical frequency domain features, SAR frequency domain features and thermal infrared frequency domain features are spliced into frequency domain feature vectors. The spatial feature vectors and frequency domain feature vectors are reduced to a unified dimension using principal component analysis. The reduced spatial feature vectors and frequency domain feature vectors are spliced using the direct splicing method to generate a comprehensive feature vector for each modality.
4. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 3, characterized in that: The mutual information value between the modalities is calculated and the fusion feature vector is calculated using a weighted sum method, including: The mutual information formula is used to calculate the mutual information value between modes, and the modal weight is generated based on the mutual information value; The weighted sum method is used to calculate the fused feature vector, and the activation function is used to linearly enhance the fused feature vector to obtain the enhanced fused feature vector.
5. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 4, characterized in that: The support vector machine (SVM) model is constructed to generate classification results, including: Collect multimodal remote sensing images with historical labels, generate historically enhanced fusion feature vectors, and construct a training set; Construct a support vector machine (SVM) model and calculate the SVM decision value of the enhanced fusion feature vector; The Platt scaling method is used to convert the SVM decision value into a probability, the maximum probability is selected as the confidence level, the corresponding category label is retained, and the classification result is generated.
6. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 5, characterized in that: Use the affine transformation method to map back to the multimodal remote sensing image coordinates, and use RGB color coding to generate visualization results, including: Use affine transformation method to map the classification results back to multimodal remote sensing image coordinates to generate interpretation masks as geospatial representations of category labels; Generate visualization results using RGB color coding for the interpretation masks.
7. The method for intelligent interpretation of remote sensing images based on multimodal fusion according to claim 6, characterized in that: The collecting of multimodal remote sensing images and preprocessing thereof include: Use remote sensing platforms to collect multimodal remote sensing images and perform preprocessing; The multimodal remote sensing images include optical, SAR and thermal infrared images; The preprocessing includes using the WGS84 coordinate system to perform georegistration on the multimodal remote sensing image, using the Retinex algorithm to perform illumination correction on the optical image, using the Lee filter to remove speckle noise from the SAR image, using the Planck function to perform radiation correction on the thermal infrared image, and linearly normalizing the thermal infrared image after radiation correction.
8. A remote sensing image intelligent interpretation system based on multimodal fusion, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The collection and extraction module is used to collect and preprocess multimodal remote sensing images, extract multimodal spatial and frequency domain features, splice the spatial and frequency domain features into a comprehensive feature vector, calculate the mutual information value between the modalities, generate modal weights, and calculate the fusion feature vector using the weighted sum method; The model interpretation module is used to build a support vector machine (SVM) model, calculate the SVM decision value, convert it into probability using the Platt scaling method, generate classification results, map it back to the multimodal remote sensing image coordinates using the affine transformation method, generate an interpretation mask, and generate visualization results using RGB color coding.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the remote sensing image intelligent interpretation method based on multimodal fusion described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the remote sensing image intelligent interpretation method based on multimodal fusion described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Man-machine collaborative remote sensing image intelligent labeling method and device based on visual large model
CN121392838A