A multi-modal remote sensing data fusion mineral prediction method

CN122761201APending Publication Date: 2026-09-15UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611140648.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

然而,上述单模态方法未能充分利用多源遥感数据的互补信息,而现有的多模态直接拼接融合方法仍属于静态融合范畴,无法实现不同模态特征的自适应动态融合,在面对多源异构遥感数据时,预测精度和泛化能力均难以满足高精度矿产预测的实际需求

Benefits of technology

1. 将多光谱、高光谱及合成孔径雷达遥感数据分别作为混合专家模型中的独立专家网络,并通过通道-空间联合注意力机制实现像元级的自适应动态加权融合;能够根据遥感影像中不同区域的地质属性自动调整各模态的贡献比例,突破了现有直接拼接或固定加权融合方式的局限,更完整地保留了空间纹理、光谱判别与构造结构三类互补矿化信息,显著提升了多源异构遥感数据的融合效率与矿化特征表达能力,有效解决了静态融合方法因无法适配区域地质差异而导致的矿化信息表征不全问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122761201A_ABST
    Figure CN122761201A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal remote sensing data fusion's mineral prediction method, comprising: obtaining multispectral, hyperspectral and SAR remote sensing data and pre-processing and space alignment, obtain standardized multi-modal data set;Parallel feature extraction is carried out to three kinds of data based on mixed expert model architecture, respectively obtain each modal feature map;Based on mixed expert model and channel-space joint attention mechanism, pixel-level adaptive dynamic weighted fusion is carried out to multi-modal feature map, and obtain primary fusion feature map;Multi-scale mineralization features are extracted by multi-level residual coding down-sampling, and encoder shallow feature map and last deep fusion feature map are obtained;Global context modeling is carried out using Transform-GCN double-branch parallel fusion architecture, and global deep feature map is obtained;Multi-level decoding up-sampling is carried out based on residual skip connection, and high-resolution mineralization feature map is restored in combination with shallow feature map;Finally, mineral target area probability map is generated and prediction result is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing data processing technology, and in particular to a mineral prediction method based on multimodal remote sensing data fusion. Background Technology

[0002] With the acceleration of the global energy transition, the demand for key metals such as copper, lithium, and nickel continues to rise, and the stability of their supply is directly related to the development of emerging industries such as new energy and energy storage.

[0003] In the mineral exploration technology system, remote sensing data, with its advantages of wide coverage, short acquisition cycle, high timeliness, and suitability for large-scale regional modeling, has become one of the core data sources for mineral exploration prediction. Among them, multispectral remote sensing data can provide surface texture and spatial morphology information related to mineralization; hyperspectral remote sensing data, with its nanometer-level spectral resolution, excels in mineral identification and spectral anomaly characterization; and synthetic aperture radar remote sensing data is highly sensitive to ground roughness, texture, and structural lines, and is adept at revealing ore-controlling structural information such as faults and folds. These three types of remote sensing data respectively contain complementary information needed for mineral exploration prediction, such as spatial texture, spectral mineral identification, and structural structure. Therefore, the organic integration of multimodal remote sensing data is a core technological direction for improving the reliability and accuracy of mineral exploration prediction.

[0004] Currently, intelligent interpretation and mineral prediction based on remote sensing imagery largely employ architectures represented by U-Net convolutional neural networks. However, existing technical solutions have several limitations. The original U-Net architecture itself only supports single-modal input and cannot directly adapt to the significant differences in imaging mechanisms, feature dimensions, and physical meanings of multispectral, hyperspectral, and synthetic aperture radar data. If uniform convolution processing or simple channel stitching post-processing is directly applied, it is easy to lose key details closely related to mineralization, making it difficult to fully represent the three types of core mineralization information. In addition, the ordinary convolutional blocks used in the original U-Net encoder are prone to gradient vanishing problems during deep network training, limiting the network depth. Meanwhile, mineralization alteration signals such as copper deposits typically account for a very low proportion in remote sensing images and exhibit significant multi-scale characteristics. Conventional convolutional structures cannot accurately extract such weak mineralization anomaly features, resulting in a high rate of missed mineralization information. Furthermore, the receptive field of a pure convolutional architecture is limited, making it unable to effectively model the spatial dependencies between distant pixels and global mineralization semantic information in large-scale remote sensing images, thus affecting the spatial continuity and geological rationality of mineral exploration prediction results. Simultaneously, in multimodal input scenarios, the skip-connection method of direct channel stitching between the encoder and decoder introduces a large number of redundant features. Since labeled samples are relatively scarce in the field of mineral exploration, this further exacerbates the risk of overfitting during model training, resulting in insufficient generalization ability. In addition, existing multimodal fusion schemes mostly adopt a static direct channel stitching method, which cannot adaptively adjust the contribution weights of each modality based on the image features of different regions. The synergistic value of multimodal data is difficult to fully realize, and there is still considerable room for improvement in the precision and accuracy of mineral exploration prediction.

[0005] In existing research, some works have attempted to apply the U-Net architecture to single-modal remote sensing mineral prediction. For example, one approach, based on the original U-Net architecture, uses Sentinel-2 multispectral satellite data as the single-modal input to construct a small-scale mining activity dataset, achieving remote sensing mapping of large-scale mining areas. Another approach proposes a spectral attention U-Net, embedding radiative transfer physics into a neural network for hyperspectral mineral segmentation tasks. Regarding multimodal fusion, some studies have employed direct channel-dimensional stitching to merge different modalities before inputting them into a U-Net for processing, and have also performed direct channel concatenation on multimodal images in remote sensing semantic segmentation tasks. However, these single-modal methods fail to fully utilize the complementary information from multi-source remote sensing data, and existing direct multimodal stitching fusion methods remain static, unable to achieve adaptive dynamic fusion of different modal features. When faced with multi-source heterogeneous remote sensing data, their prediction accuracy and generalization ability are insufficient to meet the practical needs of high-precision mineral prediction.

[0006] Therefore, there is an urgent need to propose a mineral prediction method that can effectively integrate multispectral, hyperspectral, and synthetic aperture radar remote sensing data to achieve adaptive fusion of multimodal features, accurate extraction of weak mineralization information, and efficient modeling of global geological patterns. This method would improve the reliability, precision, and accuracy of large-scale mineral target area prediction and provide core technical support for domestic mineral resource exploration. Summary of the Invention

[0007] The purpose of this invention is to provide a mineral prediction method based on multimodal remote sensing data fusion. By combining multimodal parallel feature extraction and pixel-level adaptive dynamic fusion with global context collaborative modeling and residual skip connections, the method improves the accuracy and generalization ability of mineral prediction using multi-source remote sensing data in scenarios with few samples. It also solves the technical problems of incomplete mineralization information representation, missed weak anomalies, and high risk of overfitting caused by the existing U-Net architecture due to single-modal input, static fusion, and limited receptive field.

[0008] To address the aforementioned technical problems, a first aspect of this invention provides a mineral prediction method based on multimodal remote sensing data fusion, comprising the following steps: S1: Acquire multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data of the target exploration area, perform data preprocessing and spatial alignment, and obtain a standardized multimodal dataset; S2, Parallel feature extraction is performed on the multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data in the multimodal dataset based on the hybrid expert model MoE architecture, to obtain multispectral feature maps, hyperspectral feature maps, and SAR feature maps, respectively; S3, based on the hybrid expert model MoE architecture and the channel-space joint attention mechanism, performs pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map and SAR feature map to obtain the primary fused feature map; S4, perform multi-level residual coding downsampling on the primary fusion feature map, extract multi-scale mineralization features, and obtain multiple sets of encoder shallow feature maps and final deep fusion feature maps at different scales. S5, using the Transformer-GCN dual-branch parallel fusion architecture, global context modeling is performed on the final-level deep fusion feature map to obtain a global deep feature map; S6. Based on residual skip connections, multi-level decoding upsampling is performed on the global depth feature map, and high-resolution spatial details are recovered by combining the shallow feature map of the encoder to obtain a high-resolution mineralized feature map. S7. Based on the high-resolution mineralization feature map, a mineral target area probability map is generated to obtain the mineral prediction results of the target exploration area.

[0009] Furthermore, the multi-branch parallel encoder based on the hybrid expert model MoE architecture includes: a multispectral feature extraction branch, a hyperspectral feature extraction branch, and a SAR feature extraction branch; Step S2, which involves parallel feature extraction of the multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data in the multimodal dataset to obtain multispectral feature maps, hyperspectral feature maps, and SAR feature maps, includes: S21, input the multispectral remote sensing data into the multispectral feature extraction branch, and use a lightweight 3×3 standard residual convolution architecture to extract the surface texture features and spatial morphology features of the multispectral remote sensing data to obtain a multispectral feature map. S22, the hyperspectral remote sensing data is input into the hyperspectral feature extraction branch, and a dual-branch serial structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution is used to jointly extract spectral-spatial mineralization features to obtain a hyperspectral feature map. S23, the SAR remote sensing data is input into the SAR feature extraction branch, and a multi-scale texture extraction residual convolution architecture is used to extract the ground roughness, texture features and mineral-controlling structural information of the SAR remote sensing data to obtain the SAR feature map.

[0010] Furthermore, the joint extraction of spectral-spatial mineralization features using a dual-branch serial structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution includes: S221, The hyperspectral remote sensing data is input into a multi-scale one-dimensional spectral convolution branch. The multi-scale one-dimensional spectral convolution branch includes a first one-dimensional convolution branch, a second one-dimensional convolution branch, and a third one-dimensional convolution branch arranged in parallel with convolution kernel sizes ranging from small to large. The first one-dimensional convolution branch extracts local spectral details of narrow-band mineral absorption peaks; the second one-dimensional convolution branch extracts medium-width spectral features; and the third one-dimensional convolution branch extracts wide-band spectral variation trends. S222, The output feature maps of the first one-dimensional convolution branch, the second one-dimensional convolution branch and the third one-dimensional convolution branch are concatenated by channels and then input into the next one-dimensional convolution layer. After multi-layer and multi-scale one-dimensional convolution processing, the spectral dimension encoded feature map is output. S223, The spectral dimension encoded feature map is input into the two-dimensional spatial convolution branch, and spatial dimension features are extracted through multi-layer two-dimensional convolution to achieve joint extraction of spectral-spatial mineralization features, thereby obtaining the hyperspectral feature map.

[0011] Furthermore, the method of employing a multi-scale texture extraction residual convolutional architecture to extract land cover roughness, texture features, and mineralization-controlling structural information from SAR remote sensing data includes: S231, the SAR remote sensing data is input into a multi-scale texture extraction residual convolutional architecture, the multi-scale texture extraction residual convolutional architecture includes a first convolutional layer, a second convolutional layer and a third convolutional layer connected in sequence in order from small size to large size; The first convolutional layer is used to extract fine texture features of local structural fractures. Its output is added to the input of the first convolutional layer through a residual connection and then input to the second convolutional layer. The second convolutional layer is used to extract medium-scale structural features. Its output is added to the input of the second convolutional layer through a residual connection and then input to the third convolutional layer. The third convolutional layer is used to extract macroscopic contour features of regional faults and folds. Its output is added to the input of the third convolutional layer through a residual connection. S232, after being processed step by step through the first convolutional layer, the second convolutional layer and the third convolutional layer, the SAR feature map is obtained.

[0012] Furthermore, the pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map, and SAR feature map based on the hybrid expert model MoE architecture and channel-spatial joint attention mechanism includes: S31, the multispectral feature extraction branch, hyperspectral feature extraction branch, and SAR feature extraction branch are defined as MS expert network, HSI expert network, and SAR expert network, respectively, and each outputs its own mode-specific features; S32, through learnable weights and channel-space joint attention mechanism, calculate the contribution weight of different modal features for each spatial coordinate position in the multispectral feature map, hyperspectral feature map and SAR feature map; S33, based on the regional image features, the contribution ratios of different expert networks are automatically adjusted according to the contribution weights to achieve adaptive dynamic weighted fusion of pixel-level multimodal features, thereby obtaining the primary fusion feature map.

[0013] Furthermore, the step of calculating the contribution weights of different modal features for each spatial coordinate position in the multispectral feature map, hyperspectral feature map, and SAR feature map through learnable weights and channel-space joint attention mechanism includes: S321. For each modal feature map, calculate the global statistical features of its channel dimension, and introduce the channel statistical features of the other two modalities as references. Generate the channel dimension weight vector of each modality through cross-modal interactive calculation to strengthen the mineralization-related feature channels that are complementary to other modalities in each modality. S322, the modal feature maps after being weighted by the channel dimension weight vector are fused, the response intensity of the fused feature map to each modality at each spatial location is calculated, and a spatial attention weight map for each modality is generated. S323, Multiply each modal feature map sequentially by the corresponding channel dimension weight vector and the spatial attention weight map to obtain a weighted modal feature map; S324: Input the weighted feature maps of each modality into the gated network, and calculate the normalized contribution weight of different modal features at each spatial coordinate position through the inter-modal comparison learning mechanism.

[0014] Furthermore, the step of performing multi-level residual coding downsampling on the primary fused feature map includes: S41, Construct an encoder with four coding levels, downsample the input feature map step by step, and perform a spatial resolution reduction downsampling operation on the input feature map at each coding level; S42, in each encoding level, features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units. The residual convolutional blocks use residual connection structures to encode input features and alleviate the gradient vanishing problem in deep networks. The spatial-spectral attention gating units are placed in the skip connection paths of each encoding level to weight and filter the encoder output features. The multi-scale pyramid pooling units divide the input feature map into multiple sub-regions of different scales, perform pooling operations on each sub-region, and upsample and concatenate the pooling results to extract the multi-scale characteristics of mineralization features. S43, after being processed step by step through four coding levels, outputs four sets of shallow encoder feature maps of different scales generated by the first to fourth coding levels, as well as the final deep fusion feature map output by the fourth coding level.

[0015] Furthermore, in each encoding level, the features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units, including: S4201, input the primary fused feature map or the feature map output from the previous encoding layer into the residual convolutional block of the current encoding layer. The residual convolutional block uses a residual connection structure to encode the input features and alleviate the gradient vanishing problem of deep networks, and outputs an encoded feature map. S4202, The encoded feature map is input into the spatial-spectral attention gating unit, which is set between the output of the residual convolutional block in the current encoding level and the skip connection path; S4203, in the spatial-spectral attention gating unit, the encoded feature map is reduced in dimensionality by 1×1 convolution along the spectral dimension to obtain a single-channel or low-channel spatial feature description; the spatial feature description is mapped to a spatial attention gating map by the Sigmoid activation function; the spatial attention gating map is multiplied element-wise with the encoded feature map to obtain a weighted and filtered feature map. S4204, The weighted and filtered feature map is sent to the skip connection path of the current coding level as the shallow feature map of the encoder of this level; at the same time, the unweighted coding feature map is continued to be passed to the multi-scale pyramid pooling unit. S4205, the encoded feature map is input into a multi-scale pyramid pooling unit, the input feature map is divided into multiple sub-regions of different scales and pooling operations are performed on them respectively, and the pooling results are upsampled and channel spliced ​​to extract the multi-scale characteristics of mineralization features and output the processed feature map of the current encoding level.

[0016] Furthermore, the multi-scale pyramid pooling unit divides the input feature map into multiple sub-regions of different scales and performs pooling operations on each sub-region. It then upsamples and concatenates the pooling results to extract the multi-scale characteristics of the mineralization features, including: S4231, Based on the scale statistical characteristics of mineralization and alteration anomalies in remote sensing images, the input feature map input to the multi-scale pyramid pooling unit is divided into four sub-regions of different scales. S4232, perform average pooling on each sub-region at each scale to obtain the pooling feature map at the corresponding scale; S4233 employs an inter-scale residual progressive aggregation method, which integrates small-scale pooling feature maps into large-scale pooling feature maps step by step through residual connections to form a hierarchical multi-scale mineralization feature representation, resulting in an aggregated multi-scale pooling feature map, thereby enhancing the representation ability of weak mineralization alteration signals with a very low proportion. S4234 uses bilinear interpolation that preserves mineralization features to upsample the multi-scale pooled feature map aggregated at each scale to the same spatial size as the input feature map. Spatial attention weighting is introduced during the upsampling process to obtain the upsampled multi-scale pooled feature map, so as to preserve the feature response intensity of the mineralization anomaly region. S4235 performs channel stitching on the upsampled multi-scale pooling feature maps to achieve aggregation of multi-scale mineralization context information.

[0017] Furthermore, the global context modeling of the final-level deep fused feature map, obtained by using the Transformer-GCN dual-branch parallel fusion architecture, includes: S51, Construct a Transformer-GCN dual-branch parallel fusion architecture, which includes a linear complexity Transformer sub-branch, a graph convolutional network GCN sub-branch, and a residual fusion unit; S52, the final-level deep fusion feature map is input in parallel into the linear complexity Transformer sub-branch and the graph convolutional network GCN sub-branch; S53, in the linear complexity Transformer sub-branch, the spatial dependency relationship of distant pixels in a large-scale remote sensing image is modeled through a multi-head self-attention mechanism, the spatial correlation between global mineralization semantic information and mineralization control law is extracted, and the first global feature map is output. S54, in the graph convolutional network GCN sub-branch, multimodal features are mapped to heterogeneous nodes to construct a topological graph structure of mineralization features, extract the correlation between the topological features of mineralization spatial distribution and ore-controlling structures, and output the second global feature map. S55, through the residual fusion unit, the first global feature map and the second global feature map are spliced ​​together, and the feature is optimized and fused through residual connection to output a global depth feature map that has both global spatial semantic features and mineralized topological features.

[0018] Accordingly, a second aspect of the present invention provides a mineral prediction system based on multimodal remote sensing data fusion, which predicts mineral resources in a target exploration area based on the above-mentioned multimodal remote sensing data fusion mineral prediction method, and includes the following modules: The data preprocessing module is used to acquire multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data of the target exploration area, perform data preprocessing and spatial alignment, and obtain a standardized multimodal dataset. The feature extraction module is used to perform parallel feature extraction on the multispectral remote sensing data, hyperspectral remote sensing data and SAR remote sensing data in the multimodal dataset based on the hybrid expert model MoE architecture, so as to obtain multispectral feature maps, hyperspectral feature maps and SAR feature maps respectively; The dynamic fusion module is used to perform pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map and SAR feature map based on the hybrid expert model MoE architecture and the channel-space joint attention mechanism to obtain a primary fused feature map. The residual coding module is used to perform multi-level residual coding downsampling on the primary fusion feature map, extract multi-scale mineralization features, and obtain multiple sets of encoder shallow feature maps and final deep fusion feature maps at different scales. The global modeling module is used to perform global context modeling on the final-level deep fusion feature map through the Transformer-GCN dual-branch parallel fusion architecture to obtain a global deep feature map; A multi-level decoding module is used to perform multi-level decoding upsampling on the global depth feature map based on residual skip connections, and combine it with the shallow feature map of the encoder to recover high-resolution spatial details and obtain a high-resolution mineralized feature map. The probability map generation module is used to generate a mineral target area probability map based on the high-resolution mineralization feature map, and obtain the mineral prediction results of the target exploration area.

[0019] Accordingly, a third aspect of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described multimodal remote sensing data fusion mineral prediction method.

[0020] Accordingly, a fourth aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described mineral prediction method based on multimodal remote sensing data fusion.

[0021] The above-described technical solutions of the embodiments of the present invention have the following beneficial technical effects: 1. Multispectral, hyperspectral, and synthetic aperture radar remote sensing data are used as independent expert networks in a hybrid expert model, and pixel-level adaptive dynamic weighted fusion is achieved through a channel-space joint attention mechanism. This method can automatically adjust the contribution ratio of each modality according to the geological attributes of different regions in the remote sensing image, breaking through the limitations of existing direct stitching or fixed weighted fusion methods. It more completely preserves the three complementary mineralization information of spatial texture, spectral discrimination, and tectonic structure, significantly improving the fusion efficiency and mineralization feature expression ability of multi-source heterogeneous remote sensing data. It effectively solves the problem of incomplete mineralization information representation caused by the inability of static fusion methods to adapt to regional geological differences. 2. By introducing a dual-branch parallel architecture of Transformer and graph convolutional network between the encoder and decoder, the spatial dependency relationship of distant pixels and the topological association feature of mineralization points in large-scale remote sensing images are modeled respectively. The two types of global information are synergistically complemented by residual fusion. This breaks through the inherent bottleneck of the limited receptive field of pure convolutional networks, enabling the model to capture both the spatial distribution law of ore-controlling structures and the topological structure of mineralization aggregation. This significantly improves the spatial continuity and geological rationality of mineral prediction results and solves the problem of fragmented prediction results caused by insufficient global modeling ability in existing methods. 3. By collaboratively designing residual convolutional blocks, spatial-spectral attention gating, and residual skip connections, a full-link optimization mechanism was constructed, encompassing feature encoding, attention filtering, and skip connection fusion. On the encoding side, weak mineralization-related features were enhanced and background noise was suppressed. On the skip connection side, residual addition replaced traditional channel stitching to reduce feature redundancy. This significantly improved the sensitivity for identifying low-proportion mineralization alteration signals in remote sensing images. At the same time, it effectively mitigated the risk of overfitting during model training in mineral exploration scenarios with few samples, and solved the dual technical challenges of high false negative rates and insufficient generalization ability in existing methods. Attached Figure Description

[0022] Figure 1 This is an overall flowchart of the mineral prediction method based on multimodal remote sensing data fusion provided in this embodiment of the invention; Figure 2 Here is a simplified flowchart of the mineral prediction method based on multimodal remote sensing data fusion provided in this embodiment of the invention; Figure 3 This is a schematic diagram of step S2, multimodal parallel feature extraction, provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the improved U-Net encoding and decoding process, steps S4-S6, provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of global context modeling of Transformer-GCN provided in an embodiment of the present invention; Figure 6 This is a block diagram of a mineral prediction system module based on multimodal remote sensing data fusion provided in an embodiment of the present invention.

[0023] Figure label: 1. Data preprocessing module; 2. Feature extraction module; 3. Dynamic fusion module; 4. Residual coding module; 5. Global modeling module; 6. Multi-level decoding module; 7. Probabilistic graph generation module. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0025] To facilitate understanding of the technical solutions of this invention, the relevant technical terms are explained below. Multispectral remote sensing refers to a remote sensing technology that uses multiple spectral bands for ground object detection. While the number of bands is relatively small, the data stability is high, and it excels at reflecting surface texture and spatial morphological characteristics, providing basic geological background information for mineral exploration prediction. Hyperspectral remote sensing refers to a remote sensing technology that uses continuous spectral bands with nanometer-level spectral resolution for ground object detection. It has numerous bands and continuous spectra, exhibiting high sensitivity to the diagnostic absorption characteristics of specific functional groups and metal ions in minerals, and excels at mineral identification and spectral anomaly characterization. Synthetic aperture radar is an active microwave imaging remote sensing technology that images ground objects by emitting microwave signals and receiving backscattered echoes. It is not limited by weather or lighting conditions and is highly sensitive to surface roughness, texture, and linear structures, making it suitable for extracting information on mineral-controlling structures such as faults and folds.

[0026] Please refer to Figure 1 and Figure 2 The first aspect of this invention provides a mineral prediction method based on multimodal remote sensing data fusion, comprising the following steps: S1: Acquire multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data of the target exploration area, perform data preprocessing and spatial alignment, and obtain a standardized multimodal dataset.

[0027] When conducting forecasting work in typical mineral exploration target areas such as a porphyry copper metallogenic belt in western my country, it is first necessary to acquire multispectral remote sensing data, hyperspectral remote sensing data, and synthetic aperture radar (SAR) remote sensing data for the region. Multispectral data can be obtained from medium spatial resolution images acquired by the Landsat 8 land imager, whose visible to shortwave infrared bands can reflect the basic spatial morphology of surface lithology, structure, and alteration zones. Hyperspectral data can be obtained from Hyperion spaceborne hyperspectral imaging data, whose nanometer-level spectral resolution enables precise identification of absorption peaks of characteristic minerals related to mineralization, such as hydroxyl, carbonate, and iron staining. SAR data can be obtained from Sentinel-1 C-band synthetic aperture radar images, whose active microwave imaging mechanism is highly sensitive to surface roughness, texture, and linear structures.

[0028] After acquisition, the three types of raw data require sequential preprocessing operations including radiometric correction, atmospheric correction, geometric correction, spatial registration, and numerical standardization. Radiometric correction employs radiometric calibration and the dark pixel method to eliminate sensor system errors and the influence of atmospheric path radiation. Atmospheric correction uses the FLAASH model to invert the true surface reflectance, reducing radiation distortion caused by atmospheric scattering and absorption. Geometric correction uses a quadratic polynomial model combined with a digital elevation model to control pixel positioning errors within 0.5 pixels. Spatial registration uses multispectral data as a benchmark, achieving pixel-level precise alignment between hyperspectral and SAR data through an automatic matching algorithm. Finally, min-max standardization maps the values ​​of each band to the 0-1 interval, and resamples the spatial resolution of the three types of data to a consistent 30-meter scale. The standardized multimodal dataset obtained through these processes eliminates the adverse effects of radiometric differences, geometric misalignments, and scale discrepancies on subsequent feature extraction from the data source, laying a high-quality input foundation for the collaborative utilization of multi-source remote sensing information.

[0029] In practical exploration applications, this step can be automated on a cloud-based remote sensing data processing platform or a local high-performance computing workstation. For example, for a mineralized prospective area of ​​several hundred square kilometers, the preprocessing workflow can integrate the GDAL remote sensing library and radiative transfer model interface to complete the correction and registration of dozens of images in batches. It outputs three sets of spatially aligned and numerically normalized 3D data cubes, corresponding to a unified geographic grid representation of the target area in multispectral, hyperspectral, and SAR modes, respectively. This standardized preprocessing mechanism ensures that the subsequent multi-branch feature extraction network can receive pixel-to-pixel multimodal inputs, avoiding spurious features introduced by spatial misalignment or differences in numerical distribution. This guarantees the input quality and result reliability of the mineral prediction model from the very beginning of the process.

[0030] S2, based on the hybrid expert model MoE architecture, performs parallel feature extraction on multispectral remote sensing data, hyperspectral remote sensing data and SAR remote sensing data in the multimodal dataset, and obtains multispectral feature maps, hyperspectral feature maps and SAR feature maps respectively.

[0031] For the standardized multimodal dataset output in step S1, a multi-branch parallel coding structure based on a hybrid expert model architecture is used for modality-specific feature extraction. Based on the fundamental differences in imaging physics, information representation dimensions, and mineralization response sensitivity among multispectral, hyperspectral, and SAR remote sensing data, separate feature extraction branches with distinct configurations are configured for each type, rather than using a uniform convolution kernel for indiscriminate processing. The multispectral branch employs a lightweight three-layer 3×3 two-dimensional residual convolution architecture. Through continuous residual blocks, it progressively extracts texture and morphological features such as surface lithological boundaries, ring structure morphology, and spatial distribution of alteration zones while maintaining gradient smoothness, providing regional background information for mineral prediction. The hyperspectral branch is designed as a serial dual-branch structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution. First, multiple multi-scale one-dimensional convolution operations are applied along the spectral dimension of the hyperspectral data. Different convolution kernel lengths are used to capture multi-level spectral response patterns ranging from narrow-band mineral characteristic absorption peaks to wide-band spectral trends, completing effective dimensionality reduction of high-dimensional spectral information and encoding of mineral-sensitive features. Subsequently, the encoded spectral features are fed into the two-dimensional spatial convolution branch to further aggregate mineral anomaly information in the spatial neighborhood, realizing a joint spectral-spatial mineralization feature expression. The SAR branch adopts a multi-scale texture extraction residual convolution architecture, cascading convolutional layers of three receptive field scales (3×3, 5×5, and 7×7) and supplementing them with residual connections to preserve texture details at different granularities. It is specifically used to characterize surface roughness and tectonic structure features related to mineralization, such as fault scarps, fold inflection points, and dense fracture zones, from SAR intensity and phase information. The three branches are executed completely independently and in parallel at the computation graph level, each generating a feature map with the same number of channels and spatial size, which are respectively denoted as multispectral feature map, hyperspectral feature map and SAR feature map.

[0032] This parallel feature extraction method has been validated in remote sensing prospecting practices for typical deposits such as porphyry copper deposits and altered rock gold deposits. For example, in a prediction example of a copper deposit target area in western China, the feature maps extracted by the multispectral branch effectively highlighted the intersection of ring and linear structures related to mineralization. The feature maps of the hyperspectral branch showed significant enhancement of spectral absorption peak responses in known copper mineralization outcrop areas, while the feature maps of the SAR branch clearly depicted the differences in topographic texture on both sides of the ore-controlling fault zone. By setting dedicated extraction paths for each mode, the cross-modal feature confusion and key detail obliteration problems caused by uniform convolution processing in existing technologies are avoided. The method fully preserves the three complementary core mineralization information types: spatial texture, spectral discrimination, and tectonic structure, providing a feature foundation with sufficient information and clear physical meaning for subsequent multimodal adaptive fusion.

[0033] S3, based on the hybrid expert model MoE architecture and channel-space joint attention mechanism, performs pixel-level adaptive dynamic weighted fusion of multispectral feature maps, hyperspectral feature maps and SAR feature maps to obtain a primary fused feature map.

[0034] The three modal feature maps output from step S2 are input into a fusion unit based on a hybrid expert model and a channel-spatial joint attention mechanism, performing pixel-level adaptive dynamic weighted fusion to generate a primary fused feature map. In this architecture, the multispectral feature extraction branch, hyperspectral feature extraction branch, and SAR feature extraction branch are abstracted as MS expert networks, HSI expert networks, and SAR expert networks under the hybrid expert model framework, respectively. The functional positioning of each expert network closely corresponds to its contribution to the mineralization information of the modality it processes. The MS expert is responsible for spatial texture and basic geological background representation, the HSI expert is responsible for mineral spectral identification and alteration anomaly characterization, and the SAR expert is responsible for extracting tectonic structure and surface roughness information. A learnable channel-spatial joint attention mechanism is introduced in the fusion stage to calculate the differentiated contribution weights of the three modal features at each spatial coordinate position on the multispectral, hyperspectral, and SAR feature maps. Specifically, this mechanism first adaptively weights the feature channels within each modality at the channel dimension, strengthening high-response channels related to mineralization and suppressing background noise channels. Then, at the spatial dimension, it generates spatial attention weights for each pixel location on the feature map, enabling the model to automatically adjust the contribution ratios of the three expert networks based on the geological unit type of that location, such as exposed bedrock, vegetated areas, tectonic fracture zones, or Quaternary covered areas. The fusion process does not use preset fixed weighting coefficients; instead, it learns a dynamically related weight distribution associated with the input image content through backpropagation during network training, thus achieving adaptive game theory and collaborative enhancement of multimodal features at the pixel level.

[0035] In practical mineral exploration scenarios, this dynamic fusion mechanism has clear physical significance and application value. For example, in alteration zones with well-exposed bedrock, the hyperspectral expert network receives a higher weight due to its high sensitivity to mineral absorption peaks, making the primary fused feature map highlight spectral anomalies in this area. In areas with dense tectonic fractures but indistinct spectral features, the SAR expert network contributes a higher weight due to its ability to capture texture roughness and linear structures, ensuring that tectonic mineralization information is not obscured by the multispectral or hyperspectral background. In complex land cover areas where mixed pixels are prevalent, the surface morphology and background information provided by the MS expert network serve as basal features, working in conjunction with other modal information. The resulting primary fused feature map, while retaining the unique information of each modality, achieves adaptive complementarity and spatial adaptive enhancement of cross-modal information, effectively solving key problems such as unbalanced modal contributions and poor regional adaptability in the collaborative utilization of multi-source heterogeneous remote sensing data by traditional direct channel stitching or fixed weighted fusion methods.

[0036] S4. Perform multi-level residual coding downsampling on the primary fusion feature map to extract multi-scale mineralization features, and obtain multiple sets of encoder shallow feature maps and final deep fusion feature maps at different scales.

[0037] Multi-level residual coding downsampling is performed on the primary fusion feature map generated in step S3 to extract multi-scale mineralization features and obtain multiple sets of shallow encoder feature maps and final-level deep fusion feature maps at different scales. A deep convolutional encoder with four coding layers is constructed. Each layer performs a downsampling operation that halves the spatial resolution of the input feature map. By progressively compressing the spatial dimension, the receptive field is expanded, thereby capturing multi-level mineralization features from local texture to global semantics. Within each coding layer, three key components are sequentially connected: a residual convolutional block, a spatial-spectral attention gating unit, and a multi-scale pyramid pooling unit. The residual convolutional block replaces the ordinary convolutional unit in the original U-Net. By introducing skip identity mapping connections, it effectively alleviates the gradient decay problem during backpropagation in deep networks, allowing the encoder to be stacked to deeper network layers to extract more abstract semantic features. Spatial-spectral attention gating units are positioned on the skip connection paths of each coding level. They weight and filter the output features of the current level to be passed to the decoder, generating attention weight maps based on the dependencies between feature channels and the intensity of mineralization response at spatial locations. This strengthens feature responses closely related to mineralization while suppressing irrelevant background and noise components. Multi-scale pyramid pooling units are located at the end of each coding level. They divide the input feature map into spatial sub-regions of different scales, such as 1×1, 2×2, 3×3, and 6×6, and perform pooling operations on each sub-region. After restoring the sub-regions to a uniform size through bilinear interpolation, the channels are concatenated. This aggregates contextual information from multiple receptive fields within the same level, effectively adapting to the multi-scale distribution characteristics of mineralization alteration information, ranging from regional mineralization backgrounds to local weak mineralization anomalies.

[0038] After progressive downsampling and feature refinement across four coding levels, the encoder ultimately outputs four sets of shallow feature maps with progressively decreasing sizes, corresponding to multi-scale mineralization feature representations at different spatial resolutions. Simultaneously, the fourth coding level outputs a final deep fusion feature map with increased channel count and spatial size compressed to one-sixteenth of the original input. In a specific embodiment, for multimodal input data with a 30-meter spatial resolution, the spatial sizes of the feature maps output from the first to fourth coding levels are one-half, one-quarter, one-eighth, and one-sixteenth of the original size, respectively. The shallow feature maps retain relatively rich spatial details and boundary information, helping to recover the fine contours of the prediction results during the decoding stage; the deep feature maps contain highly abstract global semantic information, reflecting large-scale ore-controlling structures and regional mineralization background features, providing condensed and high-quality feature input for subsequent global context modeling.

[0039] S5 uses the Transformer-GCN dual-branch parallel fusion architecture to perform global context modeling on the final-level deep fusion feature map, thereby obtaining a global deep feature map.

[0040] For the deep fused feature map output from the final stage of the encoder in step S4, a Transformer-GCN dual-branch parallel fusion architecture is adopted for global context modeling to generate a global deep feature map that combines global spatial semantic information and mineralization topological features. This architecture consists of two parallel feature processing sub-branches and a residual fusion unit. The first sub-branch is a linear complexity Transformer structure, with its core operation being a multi-head self-attention mechanism. This sub-branch unfolds the input deep fused feature map into a serialized feature vector and explicitly models the spatial dependencies between distant pixels in a large-scale remote sensing image by calculating the attention weights between any two spatially located features within the sequence. In geological prospecting scenarios, this global spatial modeling capability enables the model to capture the continuous extension features of ore-controlling structures along the strike direction, the gradual spatial distribution of mineralization alteration zones, and the overall spatial pattern of regional magmatic-hydrothermal activity centers and surrounding mineralization zones. This overcomes the structural constraints of pure convolutional encoders, which, despite the progressively expanding receptive field, are still limited to local neighborhoods.

[0041] The second sub-branch is a graph convolutional network structure, which maps the feature vectors of each spatial location or superpixel region in the deep fused feature map to nodes in the graph structure. It constructs a topological graph representing the correlation of mineralization features by calculating the feature similarity between nodes. Graph convolution operations aggregate feature information within the node neighborhood, thereby extracting the spatial distribution clustering patterns of mineralization points, the topological correlation between the mineralization center and the surrounding scattered mineralization points, and the feature coupling relationships presented at the intersection of different ore-controlling structures. This explicit modeling of the topological structure provides an effective means to understand the complex correlation features of the spatial distribution of mineralization in metallogenic systems. The output feature maps of the two sub-branches are then fed into the residual fusion unit, first stitched along the channel dimension to retain the independent information captured by each, and then added to the original input features through residual connections. After convolution optimization, a global deep feature map is output. This feature map simultaneously contains global semantic information based on spatial sequence dependencies and mineralization structural information based on topological graph correlations, achieving a collaborative expression of metallogenic geological laws in both spatial and topological dimensions, significantly improving the model's overall perception and understanding of the spatial structure of metallogenic systems in large-scale remote sensing images.

[0042] S6, based on residual skip connections, performs multi-level decoding upsampling on the global depth feature map, and combines it with the shallow feature map of the encoder to recover high-resolution spatial details, thus obtaining a high-resolution mineralized feature map.

[0043] A multi-level decoding upsampling structure based on residual skip connections is adopted to progressively restore the spatial resolution of the global depth feature map output in step S5. This is then fused across layers with the corresponding level encoder shallow feature map saved in step S4, ultimately yielding a high-resolution mineralized feature map with the same spatial size as the input remote sensing data. The decoder is constructed in four layers, mirroring the four-level downsampling structure of the encoder. Each decoding layer first uses bilinear interpolation to double the spatial resolution of the feature map output from the previous layer. Bilinear interpolation calculates the added pixel feature value using a weighted average of neighboring pixel values, preserving the continuity and smoothness of the feature map while restoring spatial details. The upsampled decoded features are then processed through a residual skip connection mechanism to obtain the shallow feature map of the corresponding encoder layer after spatial-spectral attention gating. This shallow feature map is first adjusted by a 1×1 convolution to make its channel count completely consistent with the channel dimension of the current decoding layer's features. Then, the two are subjected to element-wise residual addition to complete the cross-layer feature fusion. Compared to the direct concatenation of channel dimensions used in the original U-Net, the residual addition operation significantly reduces the redundant growth of feature channels while preserving shallow spatial details to supplement deep semantic features. This reduces interference from irrelevant information introduced by feature stacking in multimodal input scenarios. The fused feature map is then fed into a residual convolutional block for convolution optimization and feature refinement, before proceeding to the next decoding level for upsampling and fusion. Through four decoding levels of progressive recovery, the spatial size of the feature map gradually expands to be completely consistent with the original input remote sensing image. In the final high-resolution mineralization feature map, each pixel corresponds to a preset spatial resolution (e.g., 30 meters square) of the ground surface of the target exploration area. Its channel dimension integrates the global semantic discrimination capability of deep coding and the precise spatial positioning information transmitted by shallow skip connections at each pixel location, effectively restoring the boundary contours and internal structural details of the mineralization anomaly, resulting in higher spatial consistency between the prediction results and the actual geological morphology.

[0044] S7 generates a mineral target area probability map based on a high-resolution mineralization feature map, and obtains the mineral prediction results for the target exploration area.

[0045] Based on the high-resolution mineralization feature map output in step S6, a mineral target area probability map is generated, and the final mineral prediction result for the target exploration area is output accordingly. First, the multi-channel high-resolution mineralization feature map is input into a 1×1 convolutional layer. Through point-by-point linear combination, the multi-channel features are compressed and mapped into a single-channel feature map. The value of each pixel in this single-channel feature map comprehensively reflects the feature response intensity of the location belonging to a favorable mineralization zone. Subsequently, a sigmoid activation function is used to perform a pixel-by-pixel nonlinear transformation on the single-channel feature map, mapping the feature values ​​to the interval between 0 and 1, generating a continuously valued probability map. The probability value of each pixel represents the prediction confidence that the ground location is a mineralization target area; the closer the value is to 1, the higher the probability that the area contains mineral resources. In practical applications, probability thresholds can be set according to the specific needs of the exploration task to delineate mineral exploration target areas at different levels. For example, pixel areas with probability values ​​greater than 0.7 are designated as Class I mineral exploration target areas, indicating that the model has a high confidence in the mineral resources contained in the area, and field verification work can be prioritized. Pixel areas with probability values ​​between 0.5 and 0.7 are designated as Class II prospective mineral exploration areas, serving as reference areas for subsequent work deployment. The resulting mineral target area probability map and delineation results can be directly applied to the deployment of subsequent field geological exploration work, such as guiding the layout of geochemical sampling points, geophysical survey line planning, and trenching and drilling engineering site selection. In an application example of a porphyry copper metallogenic belt in western China, the high-probability target areas delineated using the above thresholds highly match the spatial distribution of known deposits and mineral occurrences, while also indicating several new areas with similar metallogenic geological conditions but without mineralization yet discovered, providing clear spatial guidance for subsequent exploration work deployment and demonstrating good prediction accuracy and mineral exploration guidance value.

[0046] The resulting probability maps and delineation results for mineral target areas can be directly applied to subsequent field geological exploration work, such as guiding the layout of geochemical sampling points, planning geophysical survey lines, and site selection for trenching and drilling projects. This method, based on readily available multi-source remote sensing data, uses an end-to-end deep neural network to automate the entire process from raw image input to mineral prediction output, significantly reducing the manpower and time costs required for traditional manual visual interpretation and multi-source information integration analysis. In an application example of the western porphyry copper metallogenic belt, the predicted high-probability target areas closely match the spatial distribution of known deposits and mineral occurrences, while also indicating several new sections with similar metallogenic geological conditions but without mineralization, demonstrating good prediction accuracy and mineral exploration guidance value.

[0047] This invention constructs a multimodal parallel feature extraction and pixel-level adaptive dynamic weighted fusion architecture based on a hybrid expert model. This architecture fully preserves the complementary mineralization information of multispectral, hyperspectral, and SAR remote sensing data in terms of spatial texture, spectral mineral identification, and tectonic structure. Furthermore, it effectively overcomes the inherent limitation of the receptive field of pure convolutional networks by using a global context modeling mechanism that combines Transformer and graph convolutional network in parallel branches. At the same time, the collaborative design of residual convolutional blocks, spatial-spectral attention gating, and residual skip connections achieves multi-scale enhancement of weak mineralization alteration features and simultaneous suppression of overfitting with few samples. Thus, it significantly improves the fusion efficiency of multi-source heterogeneous remote sensing data and the prediction accuracy of mineral target areas, while enhancing the generalization ability and geological interpretation rationality of the mineral exploration prediction model under large-scale complex surface conditions.

[0048] Furthermore, the multi-branch parallel encoder based on the hybrid expert model MoE architecture includes: a multispectral feature extraction branch, a hyperspectral feature extraction branch, and a SAR feature extraction branch. This encoder consists of three independent branch structures: the multispectral feature extraction branch, the hyperspectral feature extraction branch, and the SAR feature extraction branch. Each branch is designed differently to address the imaging mechanism and mineralization information characterization characteristics of its corresponding modality data.

[0049] Accordingly, such as Figure 3 As shown, step S2 involves parallel feature extraction of the multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data in the multimodal dataset, yielding multispectral feature maps, hyperspectral feature maps, and SAR feature maps, respectively, including: S21: Input the multispectral remote sensing data into the multispectral feature extraction branch, and use a lightweight 3×3 standard residual convolution architecture to extract the surface texture features and spatial morphological features of the multispectral remote sensing data to obtain a multispectral feature map.

[0050] For multispectral remote sensing data, the input is a multispectral feature extraction branch, which employs a lightweight 3×3 standard residual convolutional architecture for feature encoding. Specifically, this branch consists of three consecutive stacked 3×3 two-dimensional residual convolutional blocks, each containing three convolutional layers. Each residual convolutional block introduces an identity mapping skip connection on top of the standard convolutional layers, allowing gradient information to be directly transferred to shallower layers across convolutional layers during backpropagation. This extracts surface texture and spatial morphological features from the multispectral image while maintaining network training stability. The first residual convolutional block performs preliminary feature encoding on the input multispectral data, extracting edges, corners, and basic texture primitives from the image. The second residual convolutional block further aggregates spatial neighborhood information based on the preliminary features, forming intermediate morphological features such as lithological boundaries, linear structures, and ring structures. The third residual convolutional block further abstracts and refines the intermediate features, ultimately outputting a multispectral feature map containing the surface texture and spatial morphological features of the target area.

[0051] Multispectral data has high spatial continuity in the visible to shortwave infrared band. Its image grayscale and color changes can reflect surface geological information such as lithological boundaries, ring structures and alteration zoning. The output multispectral feature map of this branch is a characteristic expression of the above-mentioned basic geological background information.

[0052] S22. Hyperspectral remote sensing data is input into the hyperspectral feature extraction branch. A dual-branch serial structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution is used to jointly extract spectral-spatial mineralization features and obtain a hyperspectral feature map.

[0053] For hyperspectral remote sensing data, the data is input into a hyperspectral feature extraction branch, which employs a dual-branch serial structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution. Hyperspectral data is characterized by a large number of bands and high spectral resolution; its hundreds of continuous bands constitute a complete spectral curve. The electronic transitions and molecular vibrations of specific functional groups or metal ions in minerals are manifested as characteristic absorption peak positions and shapes on this curve. To address this data characteristic, the hyperspectral branch first applies multiple layers of multi-scale one-dimensional convolution operations along the spectral dimension, using convolution kernels of different lengths to perform sliding calculations on the spectral sequence, capturing details of narrow-band absorption peaks, medium-width spectral variations, and wide-band trend features, thus completing the dimensionality reduction and mineralization-sensitive spectral feature encoding of the high-dimensional spectral data. After spectral dimension encoding, the resulting features are then fed into the two-dimensional spatial convolution branch, where standard two-dimensional convolution aggregates neighborhood information in the spatial dimension, ultimately outputting a hyperspectral feature map jointly expressed by spectral and spatial dimensions.

[0054] S23. Input SAR remote sensing data into the SAR feature extraction branch, and use a multi-scale texture extraction residual convolution architecture to extract the ground roughness, texture features and mineral-controlling structural information of the SAR remote sensing data to obtain SAR feature maps.

[0055] For SAR remote sensing data, the data is input into the SAR feature extraction branch, which employs a multi-scale texture extraction residual convolutional architecture for feature encoding. Specifically, this architecture consists of three 2D convolutional layers with different receptive field sizes (3×3, 5×5, and 7×7) connected in ascending order, supplemented by residual connections to preserve texture details at different granularities. The 3×3 convolutional layer primarily extracts fine texture features such as local structural fractures and densely jointed zones; its output is added to the input via a residual connection and then passed to the 5×5 convolutional layer. The 5×5 convolutional layer further perceives medium-scale lithological interfaces and structural features based on local details; its output is also passed to the 7×7 convolutional layer via a residual connection. The 7×7 convolutional layer utilizes a larger receptive field to extract macroscopic contours and spatial distribution features such as regional fault scarps and fold flanks. Through the hierarchical convolutional processing of the three convolutional layers and the preservation of original features via inter-layer residual connections, the final output is a SAR feature map that integrates multi-scale surface roughness and structural features.

[0056] SAR remote sensing imagery is created by actively transmitting microwave signals and receiving backscattered echoes from ground features. The image's grayscale variations primarily reflect the backscattering coefficient of the ground features, and are influenced by a combination of factors, including surface roughness, differences in dielectric constant, and topographic relief. During geological prospecting, ore-controlling structures such as faults, fissures, and folds often leave distinct textures and geometric imprints on the surface. Because SAR remote sensing data is not limited by lighting or weather conditions and is highly sensitive to surface micro-topographic relief and structural linear characteristics, it can effectively capture the geometric morphology and spatial distribution information of ore-controlling structures closely related to mineralization, providing unique and crucial remote sensing data support for mineral prediction.

[0057] The three feature extraction branches mentioned above are completely parallel and do not interfere with each other at the computation graph level. Their respective output multispectral feature maps, hyperspectral feature maps, and SAR feature maps are consistent in spatial size. In the channel dimension, they respectively carry mineralization-related information that the corresponding mode is good at representing. This provides three sets of expert input features with clear functional positioning and sufficient information representation for the pixel-level adaptive dynamic fusion under the hybrid expert model framework in the subsequent step S3.

[0058] Furthermore, step S22 employs a dual-branch serial structure combining multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution to jointly extract spectral-spatial mineralization features, including: S221: Hyperspectral remote sensing data is input into a multi-scale one-dimensional spectral convolution branch. This branch includes three parallel one-dimensional convolution branches: a first one-dimensional convolution branch, a second one-dimensional convolution branch, and a third one-dimensional convolution branch. The first one-dimensional convolution branch uses a kernel size of 3 to extract local spectral details of narrow-band mineral absorption peaks. The second one-dimensional convolution branch uses a kernel size of 5 to extract medium-width spectral features. The third one-dimensional convolution branch uses a kernel size of 7 to extract wide-band spectral variation trends.

[0059] Preprocessed hyperspectral remote sensing data is input into a multi-scale one-dimensional spectral convolution branch. This branch employs a parallel three-way one-dimensional convolution structure: a first-dimensional convolution branch, a second-dimensional convolution branch, and a third-dimensional convolution branch. These three branches simultaneously perform convolution operations with different receptive fields along the spectral dimension. Specifically, the first-dimensional convolution branch uses a one-dimensional convolution operator with a kernel size of 3. Its narrow receptive field is specifically designed to capture local spectral details near narrow-band mineral absorption peaks closely related to mineralization on the hyperspectral curve, such as the sharp absorption characteristics caused by hydroxyl or carbonate group vibrations in typical altered minerals. The second-dimensional convolution branch uses a one-dimensional convolution operator with a kernel size of 5 to extract spectral variation features of moderate width, taking into account both absorption peak morphology and local spectral trend information. The third-dimensional convolution branch uses a one-dimensional convolution operator with a kernel size of 7. Its larger receptive field is used to capture overall spectral variation trends over a wide spectral range, such as the overall rise or fall of continuous spectral curves in the visible to near-infrared region. The aforementioned parallel multi-scale design enables the simultaneous capture of spectral response patterns of different widths and shapes in hyperspectral data, avoiding the scale limitations of single-size convolutional kernels in spectral feature extraction.

[0060] S222: The output feature maps of the first, second, and third one-dimensional convolutional branches are concatenated by channels and then input into the next layer of one-dimensional convolution. After multiple layers and multi-scale one-dimensional convolution processing, the spectral dimension encoded feature map is output.

[0061] The feature maps output from the first, second, and third one-dimensional convolutional branches are concatenated along the channel dimension to form a multi-channel feature tensor that integrates spectral perception results from multiple scales. This concatenated feature tensor is then input into subsequent multi-layer one-dimensional convolutional layers for progressive spectral feature encoding. Through inter-layer nonlinear transformations, higher-level spectral discrimination features are gradually abstracted. After several layers of multi-scale one-dimensional convolutional processing, a spectral dimension-encoded feature map is output. This feature map has undergone effective dimensionality reduction and feature refinement in the spectral dimension while retaining the core spectral information relevant to mineral identification.

[0062] S223: Input the spectral dimension encoded feature map into the two-dimensional spatial convolution branch, extract spatial dimension features through multi-layer two-dimensional convolution, realize the joint extraction of spectral-spatial mineralization features, and obtain the hyperspectral feature map.

[0063] The spectral-encoded feature map output from the preceding steps is input into a two-dimensional spatial convolution branch. This branch consists of multiple standard two-dimensional convolutional layers, which perform neighborhood aggregation and spatial context modeling on the spectrally encoded feature map in the spatial dimension. This fuses pixel-level spectral discrimination information with the spatial correlation features of its surrounding geological background, ultimately achieving a joint expression of spectral and spatial features. The hyperspectral feature map output after this step simultaneously contains spectral response information to the absorption characteristics of mineralization-related minerals, as well as the spatial distribution and aggregation characteristics of mineralization anomalies, providing high-quality hyperspectral modality-specific features for subsequent multimodal adaptive fusion.

[0064] In the bi-branch serial structure of the hyperspectral feature extraction branch, the multi-scale one-dimensional spectral convolution part consists of five cascaded multi-scale one-dimensional convolutional layers. The first one-dimensional convolutional layer performs feature detection on the original spectral curve. Subsequent layers, based on the encoded spectral features, perform higher-level abstraction and combination layer by layer, enabling the network to gradually extract the most relevant spectral discrimination features for mineral identification from hundreds of continuous bands. After spectral encoding, the features are fed into the two-dimensional spatial convolution branch, which consists of two two-dimensional convolutional layers. This branch aggregates neighborhood information in the spatial dimension of the spectrally encoded feature map, fusing the pixel-level spectral discrimination features with the spatial correlation of the surrounding geological background, ultimately outputting a hyperspectral feature map that combines spectral recognition capabilities with spatial positioning accuracy.

[0065] Furthermore, step S23 employs a multi-scale texture extraction residual convolutional architecture to extract land cover roughness, texture features, and mineralization-controlling structural information from SAR remote sensing data, including: S231, SAR remote sensing data is input into a multi-scale texture extraction residual convolutional architecture, which includes a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence from small to large size.

[0066] Preprocessed SAR remote sensing data is input into a multi-scale texture extraction residual convolutional architecture. This architecture consists of three two-dimensional convolutional layers with different kernel sizes, sequentially connected in ascending order of size: a 3×3 first convolutional layer, a 5×5 second convolutional layer, and a 7×7 third convolutional layer. SAR remote sensing images are formed by actively transmitting microwave signals and receiving backscattered echoes from ground features. The grayscale variations in these images primarily reflect factors such as surface roughness, dielectric constant differences, and topographic relief. Therefore, they are highly sensitive to the geometric morphology and textural features of ore-controlling structures such as tectonic lines, fault zones, and fold inflection points.

[0067] The 3×3 convolutional layer is used to extract fine texture features of local structural fractures. Its output is added to the input of the 3×3 convolutional layer through residual connections and then fed into a 5×5 convolutional layer. The 5×5 convolutional layer is used to extract medium-scale structural features. Its output is added to the input of the 5×5 convolutional layer through residual connections and then fed into a 7×7 convolutional layer. The 7×7 convolutional layer is used to extract macroscopic contour features of regional faults and folds. Its output is added to the input of the 7×7 convolutional layer through residual connections.

[0068] In this cascaded architecture, the 3×3 convolutional layer at the front has a small receptive field. Its convolution operation focuses on the details of scattering intensity changes in the local neighborhood, effectively capturing the fine texture features of local structural fractures such as dense joint and fissure zones and small fault fault surfaces. The output of this convolutional layer is added element-wise to its original input through residual connections and then used as the input for the next convolutional layer. The 5×5 convolutional layer in the middle layer has a medium-sized receptive field. It further aggregates texture information over a wider spatial range based on local details, and is used to extract surface roughness variation features related to lithological interfaces, differential weathering zones, and medium-sized structural traces. Its output is also added to its own input through residual connections and then fed into the next layer. The 7×7 convolutional layer at the end has a large receptive field, and its convolutional operations can perceive the macroscopic variation trend of backscattering intensity on a larger spatial scale. It is specifically used to extract the contour and spatial distribution features of macroscopic ore-controlling structures such as regional fault scarps, asymmetric distribution of fold flanks, and large-scale tectonic fracture zones. The output of this convolutional layer is added to its input through residual connections to complete feature enhancement at this scale. The residual connection mechanism between each convolutional layer allows gradient information to be directly backpropagated across convolutional layers during the training of multi-layer networks, effectively alleviating the gradient decay problem in deep architectures, while ensuring that texture features at different scales are not over-compressed or annihilated during the step-by-step transmission.

[0069] S232, after being processed step by step through a 3×3 first convolutional layer, a 5×5 second convolutional layer, and a 7×7 third convolutional layer, yields the SAR feature map.

[0070] After being processed through a series of 3×3, 5×5, and 7×7 convolutional layers, the multi-scale texture and structural information contained in the SAR remote sensing data, ranging from local structural fissures to regional fault folds, is extracted and aggregated layer by layer, ultimately outputting a SAR feature map. This feature map maintains the same spatial dimensions as the input data and comprehensively carries the surface roughness distribution characteristics, multi-scale texture patterns, and spatial distribution information of ore-controlling structures closely related to mineralization in the target area along the channel dimension. This provides high-quality SAR modality-specific feature representations for subsequent multimodal adaptive fusion.

[0071] Furthermore, step S3, based on the hybrid expert model MoE architecture and the channel-spatial joint attention mechanism, performs pixel-level adaptive dynamic weighted fusion of multispectral feature maps, hyperspectral feature maps, and SAR feature maps, including: S31 defines the multispectral feature extraction branch, hyperspectral feature extraction branch, and SAR feature extraction branch as MS expert network, HSI expert network, and SAR expert network, respectively, and outputs their respective modality-specific features.

[0072] The multispectral feature extraction branch, hyperspectral feature extraction branch, and SAR feature extraction branch are defined as the MS expert network, HSI expert network, and SAR expert network under the hybrid expert model architecture, respectively. Each of these three expert networks has a clearly defined functional division: the MS expert network focuses on extracting basic geological background features such as surface texture and spatial morphology from multispectral data; the HSI expert network is dedicated to capturing mineral absorption peaks and spectral anomalies related to mineralization from hyperspectral data; and the SAR expert network focuses on characterizing the spatial distribution features of surface roughness, texture structure, and ore-controlling structures from synthetic aperture radar data. After receiving input data for its corresponding mode, each expert network independently completes its mode-specific feature extraction calculations and outputs multispectral feature maps, hyperspectral feature maps, and SAR feature maps as their respective mode-specific feature representations.

[0073] S32 calculates the contribution weights of different modal features for each spatial coordinate position in the multispectral feature map, hyperspectral feature map, and SAR feature map through learnable weights and channel-space joint attention mechanism.

[0074] A learnable weighting and channel-space joint attention mechanism is introduced to calculate the contribution weight of each modal feature at each spatial coordinate position on the aforementioned three modal feature maps. This mechanism adaptively evaluates the feature channels within each modality in the channel dimension, identifying and strengthening the responses of feature channels closely related to mineralization information in each modality, while suppressing interference from redundant or noisy channels. In the spatial dimension, it performs pixel-by-pixel analysis on the differences in image content at different geographical locations on the feature maps, enabling the model to perceive the differentiated performance of the same modal feature in different geological units such as exposed bedrock areas, vegetated areas, or tectonic fracture zones. Through joint calculation in both channel and spatial dimensions, a set of quantified contribution weight values ​​is generated for each modality at each spatial coordinate position to characterize the relative importance of the modal feature at the current pixel position.

[0075] S33, based on the regional image features, automatically adjusts the contribution ratio of different expert networks according to the contribution weight, realizes the adaptive dynamic weighted fusion of pixel-level multimodal features, and obtains the primary fused feature map.

[0076] Based on the actual content of remote sensing images at different spatial locations within the target exploration area, the contribution weights calculated in the aforementioned steps are used to automatically adjust the participation ratio of the output features of the MS expert network, HSI expert network, and SAR expert network in the fusion process. For exposed altered rocks with significant spectral absorption characteristics and rich mineral identification information, the feature output of the HSI expert network receives a higher weight during fusion. For fault-developed areas with prominent tectonic linear features and dense texture information, the contribution ratio of the feature output of the SAR expert network is correspondingly increased. For areas with complex surface cover that require discrimination based on overall morphology and background information, the feature output of the MS expert network participates in the fusion as basic information. This weight allocation process is not based on preset fixed parameters, but rather iteratively optimizes the learnable weights through backpropagation during network training, ultimately forming an adaptive weight distribution that is dynamically associated with the input image content. After the above pixel-level adaptive dynamic weighted fusion, a primary fusion feature map is generated. This feature map integrates the complementary mineralization information of the three expert networks, while retaining the most discriminative modal feature responses at each spatial location, providing a sufficiently informative and spatially adaptive fusion foundation for subsequent multi-scale feature extraction by the encoder.

[0077] Furthermore, in step S32, through a learnable weights and channel-space joint attention mechanism, the contribution weights of different modal features are calculated for each spatial coordinate position in the multispectral feature map, hyperspectral feature map, and SAR feature map, including: S321. For each modal feature map, calculate the global statistical features of its channel dimension, and introduce the channel statistical features of the other two modalities as references. Generate the channel dimension weight vector of each modality through cross-modal interactive calculation to strengthen the mineralization-related feature channels that are complementary to other modalities in each modality.

[0078] For each mode in the multispectral, hyperspectral, and SAR feature maps, its global statistical features are first calculated along the channel dimension. Typically, a combination of global average pooling and global max pooling is used to obtain the descriptive vector for the channel dimension. When calculating the channel weights of a particular mode, not only is the channel statistical information of that mode used, but the channel statistical features of the other two modes are also introduced as references. Cross-modal interactive operations generate channel-dimensional weight vectors specific to each mode. This cross-modal reference mechanism allows the calculation of channel weights to perceive the complementary and redundant relationships between different modes at the feature channel level, thereby selectively strengthening mineralization-related feature channels that complement other modal information and suppressing redundant cross-modal channel responses.

[0079] S322: The modal feature maps after being weighted by the channel dimension weight vector are fused together, and the response intensity of the fused feature map to each modality at each spatial location is calculated to generate a spatial attention weight map specific to each modality.

[0080] First, each modal feature map is multiplied by its corresponding channel dimension weight vector to achieve adaptive weighting of the channel dimension. Then, the three weighted modal feature maps are fused to obtain an intermediate fused feature map. Based on this intermediate fused feature map, the response intensity of the multispectral mode, hyperspectral mode, and SAR mode at each spatial location is further calculated. By analyzing the correspondence between the spatial activation patterns of the fused feature map and the original modal feature maps, a spatial attention weight map specific to each mode is generated. This spatial attention weight map reflects the differentiated dependence of different geographical locations within the target exploration area on the three remote sensing modes. For example, the spatial response intensity of the SAR mode is higher in tectonically developed areas, while the spatial response of the hyperspectral mode is more significant in altered and exposed areas.

[0081] S323, multiply each modal feature map by the corresponding channel dimension weight vector and spatial attention weight map in turn to obtain the weighted modal feature map.

[0082] The multispectral feature map, hyperspectral feature map, and SAR feature map are sequentially multiplied by their respective channel dimension weight vectors and spatial attention weight maps to achieve joint weighting of the channel and spatial dimensions. After two element-wise multiplication operations, the key feature channels and key spatial regions related to mineralization information in the original feature maps of each modality are doubly enhanced, while irrelevant background and noise components are effectively suppressed, ultimately yielding the weighted modal feature maps for each of the three modalities.

[0083] S324: Input the weighted feature maps of each modality into the gated network, and calculate the normalized contribution weight of different modal features at each spatial coordinate position through the inter-modal comparison learning mechanism.

[0084] The weighted multispectral modal feature maps, hyperspectral modal feature maps, and SAR modal feature maps are input into a gating network. The gating network employs an inter-modal contrastive learning mechanism, calculating the relative response strength of each modal feature at the same spatial coordinate location to generate a set of normalized contribution weights. These normalized contribution weights satisfy the constraints of being non-negative and having a constant sum, directly representing the relative importance of the three modal features at each spatial coordinate location. This provides a precise quantitative basis for the subsequent step S33, which automatically adjusts the contribution ratio of each expert network based on regional image features and completes pixel-level adaptive dynamic weighted fusion.

[0085] Specifically, such as Figure 4As shown, step S4, which involves multi-level residual coding downsampling of the primary fused feature map, includes: S41, construct an encoder with four coding levels, downsample the input feature map level by level, and perform a spatial resolution reduction downsampling operation on the input feature map at each coding level.

[0086] A deep encoder with four coding levels is constructed, which receives the primary fused feature map output from step S3 as input. Each coding level performs a spatial resolution reduction downsampling operation on the feature map input to that level, typically by using convolutional or pooling layers with a stride of 2 to compress the height and width of the feature map to half the input size. Through progressive downsampling from the first to the fourth coding levels, the spatial size of the feature map decreases sequentially, while the number of channels typically increases progressively, thereby gradually expanding the semantic expressive power of the features while compressing spatial redundancy. Each of the four coding levels corresponds to a different receptive field scale. The shallow coding levels have a smaller receptive field, mainly capturing spatial details of local textures, boundaries, and small-scale mineralization alteration; the deep coding levels have progressively larger receptive fields, capable of perceiving macroscopic semantic information such as the boundaries of large-scale geological bodies, regional ore-controlling structures, and the overall spatial pattern of the metallogenic system.

[0087] In S42, at each encoding level, features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units. The residual convolutional blocks use residual connection structures to encode input features and alleviate the gradient vanishing problem in deep networks. Spatial-spectral attention gating units are placed in the skip connection paths of each encoding level to weight and filter the encoder output features. The multi-scale pyramid pooling units divide the input feature map into multiple sub-regions of different scales, perform pooling operations on each sub-region separately, and upsample and concatenate the pooling results to extract the multi-scale characteristics of mineralization features.

[0088] The feature processing flow within each encoding layer consists of three functional modules: residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units. The residual convolutional block, located at the beginning of each encoding layer, introduces identity mapping skip connections on top of standard convolution operations, allowing input features to be directly added to the convolutional output across layers. This residual connection structure provides a direct path for gradients to shallower layers during backpropagation, effectively mitigating the vanishing gradient problem common in deep network training, enabling the encoder to stack to deeper layers to extract richer abstract semantic features. The spatial-spectral attention gating unit, positioned above the skip connection path of each encoding layer, applies weighted filtering to the encoder output features before they are sent to the decoder. It generates attention weights through joint analysis of dependencies between feature channels and spatial location response strength, strengthening feature components closely related to mineralization information while suppressing irrelevant background and noise interference. The multi-scale pyramid pooling unit is located at the end of each coding level. It divides the input feature map into multiple sub-regions of different scales in space. For example, it divides the feature map into grid regions of different granularities such as 1×1, 2×2, 3×3 and 6×6. Pooling operation is performed on each sub-region to aggregate the statistical features at that scale. Then, bilinear interpolation is used to restore the pooling results of each scale to a unified spatial size and channel splicing is performed. This enables the aggregation of multi-receptive field context information within the same coding level, effectively adapting to the multi-scale distribution characteristics of mineralization and alteration information from regional mineralization background to local weak anomalies.

[0089] S43, after being processed step by step through four coding levels, outputs four sets of shallow encoder feature maps of different scales generated by the first to fourth coding levels, as well as the final deep fusion feature map output by the fourth coding level.

[0090] After the sequential processing of the four coding levels, the encoder outputs two types of feature maps. The first type consists of four shallow feature maps generated by the encoder from the first to the fourth coding levels, with their spatial dimensions decreasing sequentially. These shallow feature maps preserve multi-scale mineralization feature representations at different resolutions. These shallow feature maps are passed to the corresponding levels of the decoder via skip connections to recover the spatial details and boundary accuracy of the prediction results during upsampling. The second type is the final deep fusion feature map output after processing the residual convolution block and multi-scale pyramid pooling unit at the end of the fourth coding level. The spatial size of this feature map is compressed to one-sixteenth of the original input size, while the channel dimension contains highly condensed global mineralization semantic information after four levels of coding abstraction. This provides a high-quality feature input foundation for the global context modeling of the Transformer-GCN dual-branch architecture in the subsequent step S5.

[0091] Furthermore, in step S42, at each encoding level, the features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units, including: S4201 inputs the primary fused feature map or the feature map output from the previous encoding layer into the residual convolutional block of the current encoding layer. The residual convolutional block uses a residual connection structure to encode the input features and alleviate the gradient vanishing problem in deep networks, and outputs an encoded feature map.

[0092] The current encoding layer receives feature maps passed from the previous layer. For the first encoding layer, the input is the primary fused feature map generated in step S3; for subsequent layers, it is the processed feature map output from the previous encoding layer. This input feature map first enters the residual convolutional block of the current encoding layer. This residual convolutional block contains several standard convolutional layers and an identity mapping skip connection path that goes directly from the input to the output. After undergoing nonlinear transformations in the convolutional layers for feature encoding, the input features are added element-wise to the untransformed original input features, forming a residual connection structure. This structure provides a path for gradient information to bypass the convolutional layers and propagate directly back during backpropagation, effectively alleviating the gradient vanishing problem in deep encoders during training, enabling the network to extract richer abstract semantic features while maintaining training stability. The encoded feature map output from the residual convolutional block retains the basic information of the input features and incorporates higher-order feature representations from the convolutional transformation, and is then fed into two processing paths.

[0093] S4202, input the encoded feature map into the spatial-spectral attention gating unit, which is set between the output of the residual convolutional block in the current encoding level and the skip connection path.

[0094] The attention gating unit is physically located between the output of the residual convolutional block in the current coding level and the skip connection path. Its function is to perform weighted filtering on the features that will be passed to the decoder through the skip connection.

[0095] In S4203, within the spatial-spectral attention gating unit, the encoded feature map undergoes 1×1 convolutional dimensionality reduction along the spectral dimension to obtain a single-channel or low-channel spatial feature description. The spatial feature description is then mapped to a spatial attention gating map with values ​​in the range [0, 1] using a Sigmoid activation function. The spatial attention gating map is then element-wise multiplied with the encoded feature map to obtain a weighted and filtered feature map.

[0096] This unit first performs a 1×1 convolution operation along the spectral dimension on the input encoded feature map. Through point-by-point linear combination, the multi-channel encoded feature map is compressed and reduced in dimensionality to a single-channel or low-channel spatial feature description tensor. This tensor maintains the same spatial size as the encoded feature map, and the value of each pixel comprehensively reflects the aggregate response intensity of the multi-channel features at the corresponding location. Subsequently, a sigmoid activation function is used to perform element-wise nonlinear mapping on this spatial feature description, normalizing its numerical range to the interval between 0 and 1, generating a spatial attention gating map. Regions with values ​​close to 1 in this gating map correspond to spatial locations highly correlated with mineralization information, while regions with values ​​close to 0 correspond to locations dominated by background or noise. Finally, the spatial attention gating map is multiplied element-wise with the original encoded feature map. The gating map applies differentiated weights to each pixel of the encoded feature map in the spatial dimension, allowing the feature responses of mineralization-related regions to be preserved or even enhanced, while the feature responses of irrelevant regions are effectively suppressed, resulting in a weighted and filtered feature map.

[0097] S4204 feeds the weighted and filtered feature map into the skip connection path of the current encoding level as the shallow feature map of the encoder at that level. At the same time, the unweighted encoded feature map continues to be passed to the multi-scale pyramid pooling unit.

[0098] The weighted and filtered feature map is fed into the skip connection path of the current encoding level, serving as the shallow feature map of the encoder corresponding to that level. This shallow feature map will be fused with the features of the corresponding level in the decoder during the subsequent decoding stage to recover spatial details. Meanwhile, the original encoded feature map, unweighted by the spatial-spectral attention gating unit, continues along the main encoder path to the multi-scale pyramid pooling unit for further processing. This dual-path splitting design ensures that the skip connections deliver refined, high-quality mineralization-related features to the decoder, while the deep encoder network can still perform global semantic extraction based on complete original feature information, avoiding the risk of losing potential weak mineralization anomalies due to attention gating filtering.

[0099] S4205 inputs the encoded feature map into a multi-scale pyramid pooling unit, divides the input feature map into multiple sub-regions of different scales and performs pooling operations on them respectively, upsamples and concatenates the channels of each pooling result, extracts the multi-scale characteristics of mineralization features, and outputs the processed feature map of the current encoding level.

[0100] The unweighted encoded feature map, passed along the encoder's main path, is input to a multi-scale pyramid pooling unit. This unit first spatially divides the input feature map into multiple sub-regions of different scales, typically using four grid granularities: 1×1, 2×2, 3×3, and 6×6, corresponding to different receptive field levels from global background to local details. Average pooling is then performed on each sub-region at each scale, extracting the average response of the features within that sub-region as the statistical feature representation at that scale. Since the pooling results at different scales have different spatial dimensions, bilinear interpolation is used to upsample the pooling results at each scale to the same spatial size as the input feature map. Subsequently, the upsampled feature maps from all scales are concatenated along the channel dimension to form a comprehensive feature tensor that integrates multi-scale contextual information. This multi-scale pyramid pooling operation enables the current encoding level to simultaneously capture multi-level spatial features, from regional mineralized background fields to local weak mineralization alteration anomalies, within a single level, effectively enhancing the model's ability to represent weak mineralization signals, which constitute a very small proportion of remote sensing images. The processed feature map output by the multi-scale pyramid pooling unit is used as the final output of the current encoding level. If the current level is the first to third encoding level, the feature map will continue to be passed to the next encoding level as input. If the current level is the fourth encoding level, the feature map is the final deep fusion feature map output by the encoder, which is used for global context modeling in the subsequent step S5.

[0101] Furthermore, the multi-scale pyramid pooling unit in step S423 divides the input feature map into multiple sub-regions of different scales and performs pooling operations on each sub-region separately. It then upsamples and concatenates the pooling results to extract the multi-scale characteristics of the mineralization features, including: S4231, based on the scale statistical characteristics of mineralization and alteration anomalies in remote sensing images, the input feature map of the input multi-scale pyramid pooling unit is divided into four sub-regions of scales: 1×1, 2×2, 3×3, and 6×6.

[0102] Based on the typical scale distribution patterns of mineralization and alteration anomalies in remote sensing images, the input feature map of the multi-scale pyramid pooling unit is uniformly divided into four different granularities of sub-region grids: 1×1, 2×2, 3×3, and 6×6. The 1×1 sub-grid corresponds to the global receptive range of the entire feature map, used to capture regional mineralization background field information; the 2×2 and 3×3 sub-grids correspond to medium-scale spatial regions, which are helpful for extracting spatial structural features at the mineralization zone level; and the 6×6 sub-grid corresponds to a finer local region, specifically used to capture the detailed responses of small-scale, weak mineralization and alteration anomalies.

[0103] S4232, average pooling is performed on each sub-region at each scale to obtain the pooled feature map at the corresponding scale.

[0104] Average pooling is performed on each sub-region under the four scales described above. The pooled feature map for the corresponding scale is obtained by calculating the arithmetic mean of the feature values ​​within each sub-region. Average pooling compresses the spatial dimension while preserving the overall feature response trend within the sub-region, which helps to suppress local noise interference and enhance the robustness of feature representation. The spatial size of the pooled feature map at each scale corresponds to the number of grids in the sub-region at that scale; therefore, pooled feature maps at different scales have different spatial resolutions.

[0105] S4233 employs a progressive aggregation method based on inter-scale residuals, which integrates small-scale pooling feature maps into large-scale pooling feature maps step by step through residual connections to form a hierarchical multi-scale mineralization feature representation. This results in an aggregated multi-scale pooling feature map, which enhances the ability to represent weak mineralization alteration signals that account for a very low proportion.

[0106] A hierarchical fusion of pooling feature maps at the four scales is employed using a scale-wise residual progressive aggregation method. Specifically, following the order from smallest to largest scale, the pooling feature maps corresponding to smaller scales are upsampled to the same spatial size as the adjacent larger scale pooling feature maps. Then, they are added element-wise through residual connections, allowing the local detail information captured at the smaller scale to be progressively integrated into the macroscopic semantic representation at the larger scale. After this bottom-up, hierarchical aggregation, a set of hierarchical aggregated multi-scale pooling feature maps is formed. This residual progressive aggregation mechanism avoids the limitations of traditional pyramid pooling, where features at each scale are independent and lack information interaction. It allows macroscopic background features and local weak anomaly features to mutually reinforce each other during the aggregation process, effectively improving the model's ability to perceive and represent the extremely low proportion of weak mineralization and alteration signals in remote sensing images.

[0107] S4234 uses bilinear interpolation that preserves mineralization features to upsample the multi-scale pooled feature map aggregated at each scale to the same spatial size as the input feature map. Spatial attention weighting is introduced during the upsampling process to obtain the upsampled multi-scale pooled feature map, so as to maintain the feature response intensity of the mineralization anomaly region.

[0108] A bilinear interpolation upsampling operation that preserves mineralization features is performed on the multi-scale pooled feature map aggregated at each scale to restore its spatial resolution to the same size as the original input feature map of the pyramid pooling unit. A spatial attention weighting mechanism is introduced simultaneously during the upsampling process, calculating attention weights for each spatial location in the upsampled feature map. Regions with strong mineralization anomaly responses are assigned higher weights to maintain their feature intensity, while background regions are relatively suppressed. This approach allows the upsampled multi-scale pooled feature map to effectively maintain the feature response amplitude of mineralization anomaly regions while restoring spatial details, avoiding the feature smoothing and anomalous signal attenuation problems that may arise from upsampling interpolation operations.

[0109] S4235 performs channel stitching on the upsampled multi-scale pooling feature maps to achieve aggregation of multi-scale mineralization context information.

[0110] The upsampled multi-scale pooled feature maps at four scales are concatenated along the channel dimension to form a comprehensive feature tensor with the number of channels equal to the sum of the number of channels in each scale feature map. This channel concatenation operation integrates multi-scale mineralization context information extracted from different receptive fields in the same spatial coordinate system. This ensures that the processed feature map output from the current encoding level simultaneously contains multi-level spatial feature representations ranging from the global background field to local weak anomalies, providing an information-rich multi-scale feature foundation for further feature abstraction in subsequent encoding levels or spatial detail recovery in decoding levels.

[0111] Specifically, such as Figure 5 As shown, in step S5, the global context modeling of the final-level deep fusion feature map is performed using the Transformer-GCN dual-branch parallel fusion architecture to obtain the global deep feature map, including: S51. Construct a Transformer-GCN dual-branch parallel fusion architecture, which includes a linear complexity Transformer sub-branch, a graph convolutional network GCN sub-branch, and a residual fusion unit.

[0112] A two-branch parallel fusion architecture for global context modeling is constructed, consisting of three core components: a linear complexity Transformer sub-branch, a graph convolutional network (GCN) sub-branch, and a residual fusion unit. The linear complexity Transformer sub-branch models spatial dependencies between distant pixels in the serialized feature space, the GCN sub-branch extracts topological association patterns of mineralized features in the graph structure space, and the residual fusion unit effectively integrates the outputs of the two sub-branches. These three components together constitute a parallel processing framework for global information enhancement of the final-level deep fusion feature map from different dimensions.

[0113] S52 inputs the final-level deep fused feature maps in parallel into the linear complexity Transformer sub-branch and the graph convolutional network GCN sub-branch.

[0114] The deep fused feature map output from the final stage of the encoder in step S4 is simultaneously fed into both the linear complexity Transformer sub-branch and the Graph Convolutional Network (GCN) sub-branch for parallel processing. The spatial size of this deep fused feature map has been compressed to a smaller proportion of the original input, and the feature channels contain high-order mineralized semantic information after multi-level encoding abstraction. The two sub-branches receive the same input feature map but process it using different feature representation methods and computational logics, thereby extracting the global contextual information contained within from different perspectives.

[0115] S53, in the linear complexity Transformer sub-branch, models the spatial dependency relationship of distant pixels in a large-scale remote sensing image through a multi-head self-attention mechanism, extracts the spatial correlation between global mineralization semantic information and mineralization control law, and outputs the first global feature map.

[0116] The linear complexity Transformer sub-branch serializes the input deep fused feature map, treating each position vector in its spatial dimension as an independent sequence element. Internally, this sub-branch employs a multi-head self-attention mechanism, establishing a spatial correlation weight matrix by calculating the similarity between any two position feature vectors in the sequence. This explicitly models the spatial dependencies between distant pixels in large-scale remote sensing imagery. In geological prospecting scenarios, this global spatial modeling capability enables the sub-branch to capture the continuous extension patterns of ore-controlling structures along their strike direction, the spatial correlations between different mineralized regions within the metallogenic system, and the overall spatial pattern of regional magmatic-hydrothermal activity centers and their surrounding mineralization zones. The sub-branch ultimately outputs a first global feature map, which primarily contains global mineralization semantic information based on spatial sequence dependencies and the spatial correlation features of ore-controlling patterns.

[0117] S54, in the GCN sub-branch of the graph convolutional network, maps multimodal features to heterogeneous nodes, constructs a topological graph structure of mineralization features, extracts the correlation between the topological features of mineralization spatial distribution and ore-controlling structures, and outputs a second global feature map.

[0118] The Graph Convolutional Network (GCN) sub-branch transforms the input deep fused feature map into graph-structured data for processing. This sub-branch first maps the feature vectors corresponding to different spatial locations or superpixel regions in the feature map to nodes in the graph structure. It then establishes connections by calculating the similarity of feature expressions between nodes, thus constructing a topological graph structure reflecting the spatial distribution correlation of mineralization features. During the graph convolution operation, each node updates its own expression based on the feature information of its neighbors, allowing information to be transferred and aggregated between nodes with similar mineralization response features and potential spatial connections. This processing method effectively extracts the spatial clustering patterns of mineralization points, the topological correlation between mineralization centers and peripheral scattered mineralization points, and the feature coupling relationships at the intersection of different ore-controlling structures. This sub-branch ultimately outputs a second global feature map, which mainly contains the correlation between the spatial distribution topological features of mineralization and ore-controlling structures based on the graph topology.

[0119] S55 uses a residual fusion unit to concatenate features from the first global feature map and the second global feature map, and performs feature optimization fusion through residual connections to output a global depth feature map that combines global spatial semantic features and mineralized topological features.

[0120] The first and second global feature maps are fused using a residual fusion unit. First, the feature maps output from the two sub-branches are concatenated along the channel dimension, integrating the global spatial semantic features from the Transformer sub-branch and the mineralized topological features from the GCN sub-branch into the same tensor. Then, the concatenated features are added to the input features of the fusion unit via residual connections, followed by feature optimization through convolutional layers, ultimately outputting a global depth feature map. This residual fusion method preserves the specific types of global information extracted independently by each sub-branch and achieves synergistic complementarity between spatial sequence dependencies and topological dependencies through cross-branch feature interaction. The resulting global depth feature map possesses dual expressive power of global spatial semantic features and mineralized topological features, providing a sufficient and high-quality global contextual information foundation for the subsequent decoder to recover the high-resolution mineralized feature map.

[0121] Based on the above-described embodiments, this invention also provides several alternative implementation schemes. Regarding multi-source data expansion, on the basis of existing multispectral, hyperspectral, and SAR remote sensing data, additional dedicated feature extraction branches for newly added geophysical and geochemical exploration data can be added. A hybrid expert model framework can be used to assign corresponding expert networks to the new data modalities and incorporate them into a channel-spatial joint attention fusion mechanism, achieving unified access and collaborative fusion of more heterogeneous exploration data from multiple sources, further improving the overall accuracy of mineral prediction. Regarding the global context modeling architecture, the linear complexity Transformer sub-branch in step S5 can be replaced with lightweight global modeling architectures such as linear attention Transformer or Swing Transformer to reduce computational complexity. The graph convolutional network (GCN) sub-branch can be replaced with graph neural network architectures such as GAT or GraphSAGE to adapt to graph structure data of different scales. Regarding the multimodal fusion mechanism, the fusion method based on hybrid expert models and channel-spatial joint attention can be replaced with cross-attention fusion, inter-modal contrastive learning fusion, or cross-modal Transformer fusion to achieve deep interaction and collaborative enhancement between multimodal features. In terms of model deployment, channel pruning, parameter quantization, and knowledge distillation can be performed on each convolutional layer, Transformer layer, and fully connected layer in the trained prediction model to generate a lightweight prediction model with significantly reduced parameters and significantly improved inference speed. This enables the model to be deployed on portable computing devices or mobile terminals in the field, meeting the application needs of real-time and rapid mineral target area prediction in field exploration sites.

[0122] Accordingly, please refer to Figure 6 A second aspect of this invention provides a mineral prediction system based on multimodal remote sensing data fusion, which predicts mineral resources in a target exploration area based on the aforementioned multimodal remote sensing data fusion method, and includes the following modules: Data preprocessing module 1 is used to acquire multispectral remote sensing data, hyperspectral remote sensing data and SAR remote sensing data of the target exploration area, perform data preprocessing and spatial alignment, and obtain a standardized multimodal dataset; Feature extraction module 2 is used to perform parallel feature extraction on the multispectral remote sensing data, hyperspectral remote sensing data and SAR remote sensing data in the multimodal dataset based on the hybrid expert model MoE architecture, so as to obtain multispectral feature maps, hyperspectral feature maps and SAR feature maps respectively; Dynamic fusion module 3 is used to perform pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map and SAR feature map based on the hybrid expert model MoE architecture and channel-space joint attention mechanism to obtain a primary fused feature map; The residual coding module 4 is used to perform multi-level residual coding downsampling on the primary fusion feature map, extract multi-scale mineralization features, and obtain multiple sets of encoder shallow feature maps and final deep fusion feature maps at different scales. Global modeling module 5 is used to perform global context modeling on the final-level deep fusion feature map through the Transformer-GCN dual-branch parallel fusion architecture to obtain a global deep feature map. The multi-level decoding module 6 is used to perform multi-level decoding upsampling on the global depth feature map based on residual skip connections, and combine it with the shallow feature map of the encoder to recover high-resolution spatial details and obtain a high-resolution mineralized feature map. The probability map generation module 7 is used to generate a mineral target area probability map based on the high-resolution mineralization feature map, and obtain the mineral prediction results of the target exploration area.

[0123] Accordingly, a third aspect of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described multimodal remote sensing data fusion mineral prediction method.

[0124] Accordingly, a fourth aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described mineral prediction method based on multimodal remote sensing data fusion.

[0125] The embodiments of this invention aim to protect a mineral prediction method based on multimodal remote sensing data fusion, which has the following effects: 1. Multispectral, hyperspectral, and synthetic aperture radar remote sensing data are used as independent expert networks in a hybrid expert model, and pixel-level adaptive dynamic weighted fusion is achieved through a channel-space joint attention mechanism. This method can automatically adjust the contribution ratio of each modality according to the geological attributes of different regions in the remote sensing image, breaking through the limitations of existing direct stitching or fixed weighted fusion methods. It more completely preserves the three complementary mineralization information of spatial texture, spectral discrimination, and tectonic structure, significantly improving the fusion efficiency and mineralization feature expression ability of multi-source heterogeneous remote sensing data. It effectively solves the problem of incomplete mineralization information representation caused by the inability of static fusion methods to adapt to regional geological differences. 2. By introducing a dual-branch parallel architecture of Transformer and graph convolutional network between the encoder and decoder, the spatial dependency relationship of distant pixels and the topological association feature of mineralization points in large-scale remote sensing images are modeled respectively. The two types of global information are synergistically complemented by residual fusion. This breaks through the inherent bottleneck of the limited receptive field of pure convolutional networks, enabling the model to capture both the spatial distribution law of ore-controlling structures and the topological structure of mineralization aggregation. This significantly improves the spatial continuity and geological rationality of mineral prediction results and solves the problem of fragmented prediction results caused by insufficient global modeling ability in existing methods. 3. By collaboratively designing residual convolutional blocks, spatial-spectral attention gating, and residual skip connections, a full-link optimization mechanism was constructed, encompassing feature encoding, attention filtering, and skip connection fusion. On the encoding side, weak mineralization-related features were enhanced and background noise was suppressed. On the skip connection side, residual addition replaced traditional channel stitching to reduce feature redundancy. This significantly improved the sensitivity for identifying low-proportion mineralization alteration signals in remote sensing images. At the same time, it effectively mitigated the risk of overfitting during model training in mineral exploration scenarios with few samples, and solved the dual technical challenges of high false negative rates and insufficient generalization ability in existing methods.

[0126] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A mineral prediction method based on multimodal remote sensing data fusion, characterized in that, Includes the following steps: S1: Acquire multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data of the target exploration area, perform data preprocessing and spatial alignment, and obtain a standardized multimodal dataset; S2, Parallel feature extraction is performed on the multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data in the multimodal dataset based on the hybrid expert model MoE architecture, to obtain multispectral feature maps, hyperspectral feature maps, and SAR feature maps, respectively; S3, based on the hybrid expert model MoE architecture and the channel-space joint attention mechanism, performs pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map and SAR feature map to obtain the primary fused feature map; S4, perform multi-level residual coding downsampling on the primary fusion feature map, extract multi-scale mineralization features, and obtain multiple sets of encoder shallow feature maps and final deep fusion feature maps at different scales. S5, using the Transformer-GCN dual-branch parallel fusion architecture, global context modeling is performed on the final-level deep fusion feature map to obtain a global deep feature map; S6. Based on residual skip connections, multi-level decoding upsampling is performed on the global depth feature map, and high-resolution spatial details are recovered by combining the shallow feature map of the encoder to obtain a high-resolution mineralized feature map. S7. Based on the high-resolution mineralization feature map, a mineral target area probability map is generated to obtain the mineral prediction results of the target exploration area.

2. The mineral prediction method based on multimodal remote sensing data fusion according to claim 1, characterized in that, The multi-branch parallel encoder based on the hybrid expert model MoE architecture includes: a multispectral feature extraction branch, a hyperspectral feature extraction branch, and a SAR feature extraction branch; Step S2, which involves parallel feature extraction of the multispectral remote sensing data, hyperspectral remote sensing data, and SAR remote sensing data in the multimodal dataset to obtain multispectral feature maps, hyperspectral feature maps, and SAR feature maps, includes: S21, input the multispectral remote sensing data into the multispectral feature extraction branch, and use a lightweight 3×3 standard residual convolution architecture to extract the surface texture features and spatial morphology features of the multispectral remote sensing data to obtain a multispectral feature map. S22, the hyperspectral remote sensing data is input into the hyperspectral feature extraction branch, and a dual-branch serial structure of multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution is used to jointly extract spectral-spatial mineralization features to obtain a hyperspectral feature map. S23, the SAR remote sensing data is input into the SAR feature extraction branch, and a multi-scale texture extraction residual convolution architecture is used to extract the ground roughness, texture features and mineral-controlling structural information of the SAR remote sensing data to obtain the SAR feature map.

3. The mineral prediction method based on multimodal remote sensing data fusion according to claim 2, characterized in that, The method employs a dual-branch serial structure combining multi-scale one-dimensional spectral convolution and two-dimensional spatial convolution to jointly extract spectral-spatial mineralization features, including: S221, The hyperspectral remote sensing data is input into a multi-scale one-dimensional spectral convolution branch. The multi-scale one-dimensional spectral convolution branch includes a first one-dimensional convolution branch, a second one-dimensional convolution branch, and a third one-dimensional convolution branch arranged in parallel with convolution kernel sizes ranging from small to large. The first one-dimensional convolution branch extracts local spectral details of narrow-band mineral absorption peaks; the second one-dimensional convolution branch extracts medium-width spectral features; and the third one-dimensional convolution branch extracts wide-band spectral variation trends. S222, The output feature maps of the first one-dimensional convolution branch, the second one-dimensional convolution branch and the third one-dimensional convolution branch are concatenated by channels and then input into the next one-dimensional convolution layer. After multi-layer and multi-scale one-dimensional convolution processing, the spectral dimension encoded feature map is output. S223, The spectral dimension encoded feature map is input into the two-dimensional spatial convolution branch, and spatial dimension features are extracted through multi-layer two-dimensional convolution to achieve joint extraction of spectral-spatial mineralization features, thereby obtaining the hyperspectral feature map.

4. The mineral prediction method based on multimodal remote sensing data fusion according to claim 2, characterized in that, The method employs a multi-scale texture extraction residual convolutional architecture to extract ground roughness, texture features, and mineralization-controlling structural information from SAR remote sensing data, including: S231, the SAR remote sensing data is input into a multi-scale texture extraction residual convolutional architecture, the multi-scale texture extraction residual convolutional architecture includes a first convolutional layer, a second convolutional layer and a third convolutional layer connected in sequence in order from small size to large size; The first convolutional layer is used to extract fine texture features of local structural fractures. Its output is added to the input of the first convolutional layer through a residual connection and then input to the second convolutional layer. The second convolutional layer is used to extract medium-scale structural features. Its output is added to the input of the second convolutional layer through a residual connection and then input to the third convolutional layer. The third convolutional layer is used to extract macroscopic contour features of regional faults and folds. Its output is added to the input of the third convolutional layer through a residual connection. S232, after being processed step by step through the first convolutional layer, the second convolutional layer and the third convolutional layer, the SAR feature map is obtained.

5. The mineral prediction method based on multimodal remote sensing data fusion according to claim 2, characterized in that, The method based on the hybrid expert model MoE architecture and channel-spatial joint attention mechanism performs pixel-level adaptive dynamic weighted fusion of the multispectral feature map, hyperspectral feature map, and SAR feature map, including: S31, the multispectral feature extraction branch, hyperspectral feature extraction branch, and SAR feature extraction branch are defined as MS expert network, HSI expert network, and SAR expert network, respectively, and each outputs its own mode-specific features; S32, through learnable weights and channel-space joint attention mechanism, calculate the contribution weight of different modal features for each spatial coordinate position in the multispectral feature map, hyperspectral feature map and SAR feature map; S33, Based on the regional image features, the contribution ratios of different expert networks are automatically adjusted according to the contribution weights to achieve adaptive dynamic weighted fusion of pixel-level multimodal features, thereby obtaining the primary fusion feature map.

6. The mineral prediction method based on multimodal remote sensing data fusion according to claim 5, characterized in that, The method employs a learnable weighting and channel-spatial joint attention mechanism to calculate the contribution weights of different modal features for each spatial coordinate position in the multispectral feature map, hyperspectral feature map, and SAR feature map, including: S321. For each modal feature map, calculate the global statistical features of its channel dimension, and introduce the channel statistical features of the other two modalities as references. Generate the channel dimension weight vector of each modality through cross-modal interactive calculation to strengthen the mineralization-related feature channels that are complementary to other modalities in each modality. S322, the modal feature maps after being weighted by the channel dimension weight vector are fused, the response intensity of the fused feature map to each modality at each spatial location is calculated, and a spatial attention weight map specific to each modality is generated. S323, Multiply each modal feature map sequentially by the corresponding channel dimension weight vector and the spatial attention weight map to obtain a weighted modal feature map; S324: The weighted feature maps of each modality are input into the gated network. Through the inter-modal comparison learning mechanism, the normalized contribution weight of different modal features at each spatial coordinate position is calculated.

7. The mineral prediction method based on multimodal remote sensing data fusion according to claim 1, characterized in that, The step of performing multi-level residual coding downsampling on the primary fused feature map includes: S41, Construct an encoder with four coding levels, downsample the input feature map step by step, and perform a spatial resolution reduction downsampling operation on the input feature map at each coding level; S42, in each encoding level, features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units. The residual convolutional blocks use residual connection structures to encode input features and alleviate the gradient vanishing problem in deep networks. The spatial-spectral attention gating units are placed in the skip connection paths of each encoding level to weight and filter the encoder output features. The multi-scale pyramid pooling units divide the input feature map into multiple sub-regions of different scales, perform pooling operations on each sub-region, and upsample and concatenate the pooling results to extract the multi-scale characteristics of mineralization features. S43, after being processed step by step through four coding levels, outputs four sets of shallow encoder feature maps of different scales generated by the first to fourth coding levels, as well as the final deep fusion feature map output by the fourth coding level.

8. The mineral prediction method based on multimodal remote sensing data fusion according to claim 7, characterized in that, In each encoding level, features are processed sequentially through residual convolutional blocks, spatial-spectral attention gating units, and multi-scale pyramid pooling units, including: S4201, input the primary fused feature map or the feature map output from the previous encoding layer into the residual convolutional block of the current encoding layer. The residual convolutional block uses a residual connection structure to encode the input features and alleviate the gradient vanishing problem of deep networks, and outputs an encoded feature map. S4202, The encoded feature map is input into the spatial-spectral attention gating unit, which is set between the output of the residual convolutional block in the current encoding level and the skip connection path; S4203, in the spatial-spectral attention gating unit, the encoded feature map is reduced in dimensionality by 1×1 convolution along the spectral dimension to obtain a single-channel or low-channel spatial feature description; the spatial feature description is mapped to a spatial attention gating map by the Sigmoid activation function; the spatial attention gating map is multiplied element-wise with the encoded feature map to obtain a weighted and filtered feature map. S4204, The weighted and filtered feature map is sent to the skip connection path of the current coding level as the shallow feature map of the encoder of this level; at the same time, the unweighted coding feature map is continued to be passed to the multi-scale pyramid pooling unit. S4205, the encoded feature map is input into a multi-scale pyramid pooling unit, the input feature map is divided into multiple sub-regions of different scales and pooling operations are performed on them respectively, and the pooling results are upsampled and channel spliced ​​to extract the multi-scale characteristics of mineralization features and output the processed feature map of the current encoding level.

9. The mineral prediction method based on multimodal remote sensing data fusion according to claim 7, characterized in that, The multi-scale pyramid pooling unit divides the input feature map into multiple sub-regions of different scales and performs pooling operations on each sub-region. It then upsamples and concatenates the pooling results to extract the multi-scale characteristics of the mineralization features, including: S4231, Based on the scale statistical characteristics of mineralization and alteration anomalies in remote sensing images, the input feature map input to the multi-scale pyramid pooling unit is divided into four sub-regions of different scales. S4232, perform average pooling on each sub-region at each scale to obtain the pooling feature map at the corresponding scale; S4233 employs an inter-scale residual progressive aggregation method, which integrates small-scale pooling feature maps into large-scale pooling feature maps step by step through residual connections to form a hierarchical multi-scale mineralization feature representation, resulting in an aggregated multi-scale pooling feature map, thereby enhancing the representation ability of weak mineralization alteration signals with a very low proportion. S4234 uses bilinear interpolation that preserves mineralization features to upsample the multi-scale pooled feature map aggregated at each scale to the same spatial size as the input feature map. Spatial attention weighting is introduced during the upsampling process to obtain the upsampled multi-scale pooled feature map, so as to preserve the feature response intensity of the mineralization anomaly region. S4235 performs channel stitching on the upsampled multi-scale pooling feature maps to achieve aggregation of multi-scale mineralization context information.

10. The mineral prediction method based on multimodal remote sensing data fusion according to any one of claims 1-9, characterized in that, The Transformer-GCN dual-branch parallel fusion architecture is used to perform global context modeling on the final-level deep fusion feature map to obtain a global deep feature map, including: S51, Construct a Transformer-GCN dual-branch parallel fusion architecture, which includes: a linear complexity Transformer sub-branch, a graph convolutional network GCN sub-branch, and a residual fusion unit; S52, the final-level deep fusion feature map is input in parallel into the linear complexity Transformer sub-branch and the graph convolutional network GCN sub-branch; S53, in the linear complexity Transformer sub-branch, the spatial dependency relationship of distant pixels in a large-scale remote sensing image is modeled through a multi-head self-attention mechanism, the spatial correlation between global mineralization semantic information and mineralization control law is extracted, and the first global feature map is output. S54, in the graph convolutional network GCN sub-branch, multimodal features are mapped to heterogeneous nodes to construct a topological graph structure of mineralization features, extract the correlation between the topological features of mineralization spatial distribution and ore-controlling structures, and output the second global feature map. S55, through the residual fusion unit, the first global feature map and the second global feature map are spliced ​​together, and the feature is optimized and fused through residual connection to output a global depth feature map that has both global spatial semantic features and mineralized topological features.