Gold mine target region prediction method and system based on metallogenic causal orientation

By introducing the PCMCI algorithm and geological prior constraints into the prediction of gold ore target areas, a mineralization causal map is constructed and causal guidance features are extracted. This solves the problems of prediction accuracy and interpretability caused by the neglect of causal mechanisms in existing technologies, and achieves more efficient target area identification in shallow coverage areas.

CN122064969AActive Publication Date: 2026-05-19XIAN CENT OF GEOLOGICAL SURVEY CGS +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN CENT OF GEOLOGICAL SURVEY CGS
Filing Date
2026-04-03
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing gold target area prediction technologies ignore the causal mechanism between variables, rely on correlation and are easily affected by noise, and lack the integration of geological prior knowledge, resulting in decreased prediction accuracy and weak interpretability in shallow overburden areas.

Method used

The PCMCI algorithm, combined with geological prior constraints, is used to construct a mineralization causal map and extract causal guidance features. The target area is then predicted by optimizing the model hyperparameters through gradient descent.

Benefits of technology

This improved the robustness and interpretability of the model, enhanced the accuracy and reliability of gold target area prediction in shallow-covered areas, and provided a scientific basis to support exploration deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064969A_ABST
    Figure CN122064969A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent exploration, in particular to a gold mine target prediction method and system based on metallogenic causal orientation, which is applied to gold mine target prediction in a shallow coverage area, and comprises the following steps: acquiring earth materialization data of a to-be-predicted area, and inputting the earth materialization data into a target prediction model, obtaining a prediction result of whether the target area is a target area; wherein the target region prediction model is trained through the following steps: acquiring historical data of a target region; based on the geological prior constraint set, obtaining a historical metallogenic causal map through a PCMCI algorithm; performing causal feature extraction on the historical metallogenic causal atlas to obtain corresponding causal-oriented features; and inputting the causal-oriented features into a target region prediction model, and optimizing model hyper-parameters through a gradient descent algorithm to obtain a target region prediction model for performing gold mine target region prediction on the shallow coverage region. The method not only improves the robustness and interpretation of the model, but also realizes accurate target prediction in a shallow coverage area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent exploration technology, and in particular to a gold ore target area prediction method and a gold ore target area prediction system based on mineralization causality. Background Technology

[0002] Gold resources, as a vital mineral resource, play an irreplaceable role in technological development, financial security, high-end manufacturing, and the jewelry industry, significantly contributing to regional economic development and employment. With the increasing depletion of surface outcrops and the shrinking traditional prospecting space, shallowly covered areas (areas where bedrock is covered by thin layers of sediment, soil, or vegetation) have become key target areas for finding new deposits. While these areas present greater exploration challenges, they often possess superior mineralization geological conditions and enormous prospecting potential. Therefore, developing efficient and accurate gold deposit prediction and exploration technologies suitable for shallowly covered areas is of great significance for overcoming resource bottlenecks, expanding gold resource reserves, and ensuring resource security.

[0003] Current gold deposit target area prediction technologies typically utilize geochemical, geophysical, and remote sensing geological data as input features, establishing correlations between mineral occurrences and anomaly indicators through statistical correlation or pattern recognition. However, this approach often ignores causal mechanisms between variables, relying solely on correlations, making it susceptible to noise and confounding factors, resulting in poor prediction performance and weak interpretability. Furthermore, the lack of effective integration of prior geological knowledge makes it difficult to distinguish between primary mineralization factors and associated phenomena, leading to a significant decrease in prediction accuracy in environments with sparse data and complex interference, such as shallow overburden areas. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] In view of the above-mentioned shortcomings and deficiencies of the existing technology, this application provides a gold deposit target area prediction method and system based on mineralization causality. It solves the technical problems of directly using geological and physical data such as geochemistry, geophysics, and remote sensing as input features, and establishing the correlation between mineral occurrences and anomaly indicators through statistical correlation or pattern recognition. This approach ignores the causal mechanism between variables, relies solely on correlation, is easily interfered with by noise and confounding factors, resulting in poor prediction performance and weak interpretability. At the same time, it lacks effective integration of geological prior knowledge, makes it difficult to distinguish between the main mineralization control factors and associated phenomena, and significantly reduces prediction accuracy in environments with sparse data and complex interference, such as shallow overburden areas.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the main technical solutions adopted in this application include:

[0008] In a first aspect, embodiments of this application provide a gold deposit target area prediction method based on mineralization causality, the method being applied to gold deposit target area prediction in shallow overburden areas, the method comprising:

[0009] Obtain geophysical data of the region to be predicted, and input the geophysical data into a pre-trained target area prediction model to obtain the prediction result of whether the target area is the target area;

[0010] The target area prediction model is trained through the following steps:

[0011] Acquire historical data for the target area, including historical geophysical data and multiple known mineral deposits within the target area, with each known mineral deposit having a corresponding known mineral deposit label;

[0012] Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data and obtain the corresponding historical mineralization causal map.

[0013] Causal features were extracted from historical mineralization causal maps to obtain corresponding causal guidance features;

[0014] The causal guidance features are input into the target area prediction model, and the model hyperparameters are optimized by the gradient descent algorithm to obtain a target area prediction model for gold mine target area prediction in shallow-covered areas.

[0015] Optionally, in one specific embodiment, the historical geophysical data includes all geochemical data, geological structure data, topographic data, and environmental data within the target area;

[0016] Based on a pre-defined set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data, obtaining corresponding historical mineralization causal maps, including:

[0017] Match all geochemical data, geological structure data, topographic data, and environmental data within the target area with all mineral deposit labels;

[0018] Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structure data, topographic data and environmental data corresponding to each mineral deposit label, and the causal path of mineralization between the mineral deposit label and the label.

[0019] Based on the mineralization causal path corresponding to each mineralization point label, a corresponding historical mineralization causal map is obtained.

[0020] Optionally, in a specific embodiment, each geochemical data, geological structure data, topographic data, and environmental data has corresponding sampling coordinates; each mineral point label has corresponding mineral point coordinates;

[0021] The geochemical data, geological structure data, topographic data, and environmental data within the target area are matched with all mineral deposit labels, including:

[0022] Based on the sampling coordinates corresponding to each geochemical data, geological structure data, topographic data, and environmental data within the target area, and using the sampling coordinates of each geochemical data as the aggregation center, all geological structure data, topographic data, and environmental data within a preset range are aggregated to obtain multiple aggregated datasets within the target area. During the aggregation process, when multiple data of the same type appear, sampling is performed based on a predefined value selection strategy.

[0023] The value selection strategy is as follows: when the data type of the same type of data is numerical data, take the arithmetic mean of all data in the same type of data; when the data type of the same type of data is categorical / coded data, take the data with the highest proportion in the same type of data.

[0024] Verify the integrity of each aggregated dataset and fill in missing data based on pre-set default values;

[0025] Based on the fact that each aggregated dataset contains geochemical data with corresponding sampling coordinates and each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels.

[0026] Optionally, in a specific embodiment, based on the sampling coordinates corresponding to the geochemical data in each aggregated dataset and the fact that each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels, including:

[0027] Based on the sampling coordinates corresponding to the geochemical data in each aggregated dataset and the corresponding mineral point coordinates for each mineral point label, and using a pre-set formula (Formula 1), the distance between each aggregated dataset and each mineral point label is obtained; Formula 1 is:

[0028] D n,m =R×arccos[sinB n sinB m +cosB n cosB m cos(L n -L m )];

[0029] Where n is the index of the aggregated dataset, m is the index of the mining point label, (Bn ,L n ) represents the sampling coordinates of the aggregated dataset with index n, (B m ,L m ) represents the coordinates of the mineral point labeled with index m, R is the Earth's radius, and D is the Earth's coordinates. n,m The distance between the aggregate dataset with index n and the mining point labels with index m;

[0030] Based on the distance between each aggregated dataset and each mining point label and the pre-set proximity principle, all aggregated datasets within the target area are matched with all mining point labels.

[0031] Optionally, in a specific embodiment, based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structural data, topographic data, and environmental data corresponding to each mineral deposit label, and the ore-forming causal path between these data and the mineral deposit label, including:

[0032] Based on the PCMCI algorithm, each geochemical data, geological structure data, topographic data and environmental data in each aggregated dataset is transformed into corresponding map nodes;

[0033] Based on each mineral point label, all map nodes in each aggregated dataset, and a pre-set set of geological prior constraints, the mineralization causal path between each aggregated dataset and the corresponding mineral point label is obtained.

[0034] Optionally, in a specific embodiment, based on each mineralization point label, all map nodes in each aggregated dataset, and a pre-set set of geological prior constraints, the mineralization causal path between each aggregated dataset and the corresponding mineralization point label is obtained, including:

[0035] Based on all the graph nodes in each aggregated dataset and a pre-set set of geological prior constraints, the causal value between every two graph nodes is obtained. Based on the causal value between every two graph nodes and a pre-set causal threshold, a corresponding graph edge is generated between two graph nodes whose causal value is greater than the causal threshold. Each graph edge has a corresponding causal direction and causal strength.

[0036] Based on the causal direction corresponding to each graph edge in each aggregated dataset, all the top-level and bottom-level graph nodes in each aggregated dataset are selected, and graph edges are formed between each top-level graph node and the corresponding mineral point label in the aggregated dataset; among them, the connection relationship between the bottom-level graph node and the corresponding mineral point label in each aggregated dataset is the mineralization causal path.

[0037] Optionally, in one specific embodiment, the geological prior constraint set is obtained through the following steps:

[0038] Object extraction is performed on the pre-set geological mineralization expert theory to obtain multiple variable objects, and all variable objects are transformed into recognizable objects by the PCMCI algorithm;

[0039] Based on pre-set constraint rules, rule constraints are applied to all identifiable objects to obtain a set of geological prior constraints; wherein, the rule constraints include: causal direction constraints, conditional independence constraints, and variable priority constraints.

[0040] Optionally, in a specific embodiment, the variable object includes a numerical geological variable object or a categorical / qualitative geological variable object;

[0041] Convert all variable objects into objects recognizable by the PCMCI algorithm, including:

[0042] Based on each numerical geological variable object and the pre-set quantitative indicators for each numerical geological variable object, each numerical geological variable object is quantitatively processed.

[0043] Based on each category / qualitative geological variable object and a pre-set coding strategy, obtain the coded value corresponding to each category / qualitative geological variable object;

[0044] The object format is unified for each quantified numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object to obtain the recognizable object of the PCMCI algorithm for each variable object;

[0045] The unified object format is as follows: each quantized numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object are stored in a structured form to adapt to the input format of the PCMCI algorithm.

[0046] Optionally, in one specific embodiment, the historical mineralization causal map includes multiple mineralization causal paths;

[0047] Causal features were extracted from historical mineralization causal maps to obtain corresponding causal guidance features, including:

[0048] Feature extraction, feature integration, and standardization are performed on all causal paths corresponding to each mineralization point label in the historical mineralization causal map to obtain causal guidance features corresponding to each mineralization point label in the historical mineralization causal map; the causal guidance features include: intensity weighted features, causal interaction features, causal contribution features, and causal existence features.

[0049] Secondly, embodiments of this application provide a gold ore target area prediction system based on mineralization causality, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the aforementioned gold ore target area prediction method based on mineralization causality.

[0050] (III) Beneficial Effects

[0051] This application presents a gold deposit target area prediction method based on mineralization causality guidance. By introducing the PCMCI algorithm and combining it with geological prior constraints to construct a mineralization causality map and extract geologically significant causal guidance features, it not only improves the robustness and interpretability of the model, but also more effectively supports accurate target area prediction in the challenging scenario of shallow cover areas. Attached Figure Description

[0052] Figure 1 A flowchart illustrating the target area prediction model training process provided in this application embodiment;

[0053] Figure 2 A flowchart for constructing historical mineralization causal maps provided in this application embodiment;

[0054] Figure 3 A flowchart illustrating the construction of the geological prior constraint set provided in this application embodiment. Detailed Implementation

[0055] To better explain and facilitate understanding of this application, the following detailed description of the application is provided in conjunction with the accompanying drawings and specific embodiments.

[0056] Gold mines, as a key strategic mineral, play a prominent role in financial security, industrial development, and regional economic growth. With surface outcrops gradually depleting, shallowly overburdened areas have become the core target areas for finding new deposits. These areas possess excellent mineralization conditions but face significant exploration challenges. Therefore, developing efficient and accurate gold deposit prediction and exploration technologies suitable for these regions is of great significance for overcoming resource bottlenecks and ensuring resource security.

[0057] Existing gold deposit target area prediction technologies mostly rely directly on geological and physical data, establishing correlations through statistical correlation or pattern recognition, but neglecting the causal mechanisms between variables and lacking the integration of geological prior knowledge. These technologies are susceptible to interference, have weak interpretability, and suffer from insufficient accuracy in scenarios with sparse data and complex interference in shallow-covered areas. This application introduces the PCMCI algorithm, combining it with geological prior constraints to construct a mineralization causal map and extract geologically significant causal guidance features. This improves the model's robustness and interpretability, and provides effective support for accurate target area prediction in shallow-covered areas.

[0058] To better understand the above technical solutions, exemplary embodiments of this application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application can be understood more clearly and thoroughly, and that the scope of this application can be fully conveyed to those skilled in the art.

[0059] This application provides a gold deposit target area prediction method based on mineralization causality. The method is applied to gold deposit target area prediction in shallow overburden areas, including:

[0060] Obtain geophysical data of the region to be predicted, and input the geophysical data into a pre-trained target area prediction model to obtain the prediction result of whether the target area is the target area;

[0061] Among them, such as Figure 1 As shown, the target area prediction model is trained through the following steps:

[0062] S10. Obtain historical data of the target area, including historical geophysical data of the target area and multiple known mineral deposits in the target area, and each known mineral deposit has a corresponding known mineral deposit label;

[0063] S20. Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data to obtain the corresponding historical mineralization causal map.

[0064] S30. Extract causal features from historical mineralization causal maps to obtain corresponding causal guidance features;

[0065] S40. Input the causal-guided features into the target area prediction model, optimize the model hyperparameters through the gradient descent algorithm, and obtain the target area prediction model for gold mine target area prediction in shallow-covered areas.

[0066] This embodiment provides a gold deposit target area prediction method based on mineralization causality guidance. It introduces a target area prediction model and uses the PCMCI algorithm to mine the causal relationship between geophysical data and mineralization under geological constraints. It constructs a historical mineralization causal map and extracts causal guidance features, effectively overcoming the shortcomings of traditional methods that rely on statistical correlation, are susceptible to noise interference, have weak generalization ability, and lack clear geological significance. The trained target area prediction model can not only more accurately identify potential gold deposit target areas, but also enhance the interpretability of the results by visualizing causal paths, providing a scientific basis for exploration deployment and significantly improving the efficiency and success rate of mineral exploration in shallow cover areas.

[0067] Optionally, in one specific embodiment, geophysical data and historical geophysical data include all geochemical data, geological structure data, topographic data and environmental data within the target area.

[0068] Specifically, acquiring geophysical data for the target area needs to cover the geometric dimensions of "source of ore-forming materials, migration channels, occurrence environment, and dynamic regulation" to ensure that the data can support subsequent causal path mining (such as "deep Au source" → fault-guided ore → surface anomaly).

[0069] Geochemical data, used to directly indicate the distribution and migration of ore-forming elements, typically comes from borehole core testing (LA-ICP-MS), selective chemical extraction of soil, and in-situ micro-area analysis. The format includes CSV (elemental content) and LAS (well logging curves). Core indicators include deep Au / As / Sb content, surface active element content, and nano gold migration index. Further subdivided into: deep geochemical data, including but not limited to the content and grade of major elements (SiO2, Al2O3), ore-forming elements (Au, As, Sb, Hg) in borehole cores, nano-gold migration index (active nano-gold content / total Au content), and correlation coefficients of ore-forming elements (such as Au-As correlation coefficient); surface geochemical data: total amount and active state content (such as ionic state, organically bound state) of ore-forming elements (Au, As, Sb) in soil / stream sediments, elemental anomaly intensity (elemental content in anomaly areas / background value), and elemental zoning characteristics (such as vertical zoning and horizontal zoning parameters); and derived geochemical data: ore-forming element migration efficiency (deep elemental content / surface elemental content), and elemental combination anomaly index (Au+As+Sb comprehensive anomaly score).

[0070] Geochemical data is typically acquired in the following ways: geochemical data in historical geophysical data is obtained by collecting geological survey reports and geochemical survey / detailed survey data (such as 1:50,000 soil geochemical measurement data) of the target area; geochemical data in geophysical data (in the field) is obtained through field sampling and testing, that is, for shallow overburden areas, borehole cores (depth 50-200m), surface soil (0-20cm), and stream sediment samples are collected, and the elemental content is tested by ICP-MS (inductively coupled plasma mass spectrometry) to calculate the proportion of active states and the intensity of anomalies.

[0071] Geological structural data is used to control the migration paths and enrichment conditions of ore-forming elements. Its main sources are 1:25,000 geological maps (SHP format), borehole imaging logging (JSON format), and seismic exploration data. Core indicators include fault distribution, fractured section thickness, and contact zone distance of granite bodies. Further subdivided into: fault-related data, namely fault distribution vector data (latitude and longitude coordinates, strike, dip, and dip angle), fault density (total fault length / area within the target area), tectonic ore-conducting capacity index (fault porosity × fracture zone thickness / 100); distance between faults and sampling points (e.g., straight-line distance from the convergence center to the nearest fault); rock mass-related data, namely the distribution range of granite and volcanic rock masses (vector boundary), rock mass contact zone distance (distance from the convergence center to the rock mass contact zone), and rock mass mineralization specific parameters (e.g., SiO2 content and rare earth element distribution pattern of granite); and stratigraphic-related data, namely stratigraphic lithology type (e.g., sandstone, granite, shale, coded), stratigraphic thickness (average thickness of the target strata within the convergence range), and stratigraphic mineralization grade (labeled as "high / medium / low" based on historical data, coded as 1 / 2 / 3).

[0072] The acquisition of geological structural data is usually carried out as follows: geochemical data in historical geophysical data is obtained by collecting geological maps (1:50,000 / 1:100,000, etc.), structural outline maps, and rock mass survey reports of the target area, and extracting the vector coordinates and attributes of faults, rock masses, and strata through ArcGIS; geochemical data in geophysical data (field verification) is obtained by combining data such as fault occurrence and rock mass contact zone location supplemented by geological mapping and borehole logging with geochemical data in historical geophysical data.

[0073] Topographic and geomorphological data are used to influence the surface migration efficiency of ore-forming elements and the preservation of anomalies. The sources are 30m resolution DEM (TIFF format) and borehole exposure data (CSV format). The core indicators include overburden thickness, elevation, slope, and aspect.

[0074] Topographic data typically includes: basic topographic data, namely, elevation, average slope, slope variability, and aspect (e.g., sunny slope, shady slope, coded) extracted from a digital elevation model (DEM, resolution ≥ 30m); overburden-related data, namely, overburden thickness (average thickness of overburden exposed by boreholes within the aggregation range, in meters), and overburden lithology (e.g., clay layer, sand layer, gravel layer, coded); and geomorphic unit data, namely, geomorphic type (e.g., mountain, valley, alluvial fan, coded), and elevation gradient (elevation difference between the aggregation center and a 1km radius around it).

[0075] Topographic data acquisition methods are generally divided into: public data, downloading DEM data (such as ASTER GDEM) from geospatial data cloud and extracting topographic parameters through ArcGIS; field and historical data, collecting borehole columnar sections to obtain overburden thickness and lithology, and combining geological maps to divide geomorphic units.

[0076] Environmental data is used to regulate the dissolution and migration efficiency of ore-forming elements. Its sources typically include: 1. quarterly sampling from groundwater monitoring points (CSV format), Landsat 8 remote sensing inversion (TIFF format), and meteorological station data (CSV format). Core indicators include groundwater pH, vegetation NDVI, and total precipitation.

[0077] Environmental data typically include: hydrological environmental data, namely groundwater pH (quarterly average), groundwater salinity (mg / L), groundwater migration contribution coefficient (Au content in groundwater samples / Au content in deep cores), and surface water runoff (annual average, in m³). 3 / s); vegetation environment data, namely vegetation cover (seasonal average NDVI), vegetation type (e.g., coniferous forest, broad-leaved forest, herbaceous, coded), and adsorption coefficient of vegetation for elements (calculated based on vegetation sample test data); climate environment data, namely annual average precipitation (mm), annual average temperature (°C); seasonal distribution coefficient of precipitation (rainy season precipitation / total annual precipitation).

[0078] Environmental data acquisition methods typically include: public data, such as obtaining temperature and precipitation data from meteorological department websites and retrieving NDVI mean values ​​from MODIS satellite data; field sampling and testing, such as collecting groundwater and surface water samples to test pH values ​​and mineralization, collecting vegetation samples to test elemental content, and calculating adsorption coefficients; and historical data, such as collecting hydrogeological survey reports and environmental monitoring data.

[0079] This embodiment integrates multi-source geophysical data covering four dimensions: "source of ore-forming materials, migration channels, occurrence environment, and dynamic regulation" (including depth and surface geochemistry, fine geological structures, high-resolution topography and geomorphology, and dynamic environmental data). By accurately depicting hidden mineral clues through key parameters such as overburden thickness, active elements, and ore-guiding structures, it effectively overcomes the bottleneck of traditional methods in this type of area, which "can see anomalies but cannot accurately determine their causes," and provides solid technical support for the efficient deployment of exploration projects and the expansion of gold resource reserves.

[0080] Optionally, in a specific embodiment, based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data to obtain corresponding historical mineralization causal maps, such as... Figure 2 As shown, it includes:

[0081] S21. Match all geochemical data, geological structure data, topographic data, and environmental data within the target area with all mineral deposit labels;

[0082] S22. Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structure data, topographic data and environmental data corresponding to each mineral deposit label, and the causal path of mineralization between the mineral deposit label and the mineral deposit label.

[0083] S23. Based on the mineralization causal path corresponding to each mineralization point label, obtain the corresponding historical mineralization causal map.

[0084] This embodiment precisely matches geochemical data, geological structural data, topographic data, and environmental data with mineral deposit labels. Based on a pre-set set of geological prior constraints, it uses the PCMCI algorithm to mine the causal paths of mineralization corresponding to each mineral deposit label, ultimately constructing a historical mineralization causal map. This method achieves semantic alignment between data and mineral deposits, eliminates spurious associations that do not conform to geological laws, and reveals the differentiated mineralization control mechanisms of different types of deposits. It not only improves the accuracy and reliability of mineral exploration prediction but also provides scientific basis and visualization support for understanding regional metallogenic laws, optimizing mineral exploration models, and formulating exploration strategies.

[0085] Furthermore, data acquisition needs to meet the following requirements: coordinate consistency, spatial resolution, temporal scale consistency, and data integrity.

[0086] Coordinate consistency requirement: All data sampling / extraction coordinates must be unified to the 2000 National Geodetic Coordinate System (such as geodetic coordinates B / L / H or plane coordinates X / Y / Z) to avoid projection deviations affecting subsequent spatial aggregation and matching.

[0087] Spatial resolution requirements: Point data (such as soil sampling, borehole data), sampling density ≥ 1 point / 500m (mineralization signals are scattered in shallow cover areas, and sufficient density is required to support aggregation); Raster data (such as DEM, NDVI), resolution ≥ 30m (adapting to an aggregation range of 30-50m to ensure the spatial correlation of the aggregated data).

[0088] Consistency of time scale: Dynamic environmental data (such as pH value, NDVI) should be unified to the annual average or the average of the critical period of mineralization (such as the rainy season) to avoid aggregation distortion caused by time differences.

[0089] Data integrity requirements: The missing rate of core variables (such as deep Au content, fault density, overburden thickness, and groundwater pH) in a single aggregated dataset should be ≤10%, and the missing rate of non-core variables (such as vegetation adsorption coefficient) should be ≤30%, to ensure the effectiveness of subsequent causal mining.

[0090] The priority for data acquisition is as follows: First priority (core mineralization factors), deep Au / As / Sb content, tectonic ore-conducting capacity index, overburden thickness, and surface active element anomaly intensity.

[0091] Second priority (synergistic mineralization factors): groundwater pH value, fault density, rock mass contact zone distance, and mean NDVI value;

[0092] The third priority (auxiliary constraint factors) includes stratigraphic lithology, annual precipitation, slope, and vegetation type.

[0093] The geophysical data obtained through the above classification can fully cover the entire chain of gold mineralization in shallowly covered areas, from "ore supply to ore guidance to migration to enrichment," providing high-quality basic data support for subsequent data aggregation, spatial matching, construction of historical mineralization causal maps, and model training.

[0094] Furthermore, when acquiring all geochemical data, geological structure data, topographic data, and environmental data within the target area, data preprocessing is required.

[0095] For example:

[0096] The geochemical data processing procedure is as follows: Elemental content sequences at depths of 50-100m are extracted from LAS logging data and merged with surface soil elemental data in CSV format to generate a "depth-elemental content" comparison table; negative values ​​caused by instrument errors and high surface elemental values ​​caused by human pollution (e.g., surface Au content at a certain point is 10 times the background value but there is no deep data to support them) are identified using the "3σ principle" and directly discarded; "Nano gold migration index = deep nano gold content / surface nano gold content" and "elemental anomaly intensity = (measured value - background value) / background value" are calculated to supplement mineralization-related variables; finally, standardization is performed by Z-score standardization of all elemental contents and derived indices to eliminate dimensional differences.

[0097] The process of processing geological structural data: GeoPandas is used to read SHP format geological maps and extract the strike, length, and density (number of faults / km). 2 Extract fault porosity and fractured section thickness from imaging logging data and convert them to CSV format; calculate "structural ore-conducting capacity index = fractured section thickness (m) × fault porosity (%) / 100" to quantify fault ore-conducting potential; associate structural data with other data according to "sampling point coordinates" (such as fault density within 1km of a sampling point, straight-line distance from the contact zone of the granite body); for areas without imaging logging data, fill in "average fault porosity of the same structural zone".

[0098] The topographic data processing procedure is as follows: The elevation, slope (°), and aspect (encoded as 0-360°) of each sampling point are extracted from the TIFF format DEM using the GDAL tool and converted into CSV format; the overburden thickness data (measured value) revealed by boreholes is used to calibrate the overburden thickness (predicted value) retrieved from the DEM and correct errors (such as measured values ​​±20% of the retrieved DEM value); the elevation gradient is calculated as "(sampling point elevation - average elevation of 1km surrounding area) / 1km" to characterize the influence of topographic relief on element migration; and the slope, elevation gradient, and overburden thickness are standardized using Z-score.

[0099] The environmental data processing procedure is as follows: the quarterly NDVI mean of each sampling point is extracted from the TIFF format remote sensing data and merged with the groundwater pH value and total precipitation data into a "sampling point-time-environmental index" matrix; a unified time scale is set according to "quarter" (e.g., the groundwater pH value is the quarterly mean and the precipitation value is the quarterly total) to ensure that it matches the time cycle of ore-forming element migration; the "groundwater migration contribution coefficient = water sample Au content / deep core Au content" is calculated to quantify the degree of groundwater's promotion of element migration; extreme precipitation values ​​(e.g., precipitation data from the season of extreme rainstorms) and abnormal groundwater pH values ​​(<4 or > 9, exceeding the stable migration range of ore-forming elements) are removed.

[0100] After the above data preprocessing, a sample-variable structured matrix (usually in CSV format) is generated, with each row corresponding to a sampling location (i.e., a sample) and each column corresponding to a quantitative variable (including core indicators and derived variables of 4 types of data).

[0101] This embodiment effectively addresses the inconsistencies in format, scale, quality, and semantics among multi-source heterogeneous geophysical data through a data preprocessing workflow, significantly improving the reliability and accuracy of subsequent causal modeling and target area prediction. By integrating deep-surface data, intelligently removing outliers, and constructing mineralization-related derived variables (such as the nano-gold migration index and tectonic mineralization capacity index), key geological information is preserved, and the ability to characterize mineralization processes is enhanced. Spatial alignment, temporal scale unification, and appropriate filling of missing values ​​ensure the integrity and representativeness of the samples. The resulting standardized structured matrix provides high-quality, comparable, and geologically significant input features for PCMCI causal discovery and machine learning models. This preprocessing method not only reduces the impact of noise and human interference but also lays a solid data foundation for the accurate identification of gold ore targets in shallow-covered areas.

[0102] Furthermore, in a specific embodiment, each geochemical data, geological structure data, topographic data, and environmental data has corresponding sampling coordinates; each mineral point label has corresponding mineral point coordinates.

[0103] Then, geochemical data, geological structure data, topographic data, and environmental data within the target area are matched with all mineral deposit labels, including:

[0104] Based on the sampling coordinates corresponding to each geochemical data, geological structure data, topographic data, and environmental data within the target area, and using the sampling coordinates of each geochemical data as the aggregation center, all geological structure data, topographic data, and environmental data within a preset range are aggregated to obtain multiple aggregated datasets within the target area. During the aggregation process, when multiple data of the same type appear, sampling is performed based on a predefined value selection strategy.

[0105] The value selection strategy is as follows: when the data type of the same type of data is numerical, take the arithmetic mean of all data in the same type of data; when the data type of the same type of data is categorical / coded data, take the data with the highest proportion in the same type of data.

[0106] Verify the integrity of each aggregated dataset and fill in missing data based on pre-set default values;

[0107] Based on the fact that each aggregated dataset contains geochemical data with corresponding sampling coordinates and each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels.

[0108] Furthermore, based on the fact that each aggregated dataset contains geochemical data with corresponding sampling coordinates and each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels, including:

[0109] Based on the sampling coordinates corresponding to the geochemical data in each aggregated dataset and the corresponding mineral point coordinates for each mineral point label, and using a pre-defined formula (Formula 1), the distance between each aggregated dataset and each mineral point label is obtained; Formula 1 is:

[0110] D n,m =R×arccos[sinB n sinB m +cosB n cosB m cos(L n -L m )];

[0111] Where n is the index of the aggregated dataset, m is the index of the mining point label, (B n ,L n ) represents the sampling coordinates of the aggregated dataset with index n, (B m ,L m ) represents the coordinates of the mineral point labeled with index m, R is the Earth's radius, and D is the Earth's coordinates. n,mThe distance between the aggregate dataset with index n and the mining point labels with index m;

[0112] Based on the distance between each aggregated dataset and each mining point label and the pre-set proximity principle, all aggregated datasets within the target area are matched with all mining point labels.

[0113] Specifically, clearly define the aggregation center and the sampling coordinates for each geochemical data point (unified in the 2000 National Geodetic Coordinate System, in latitude and longitude (B,L) or plane coordinates (X,Y)).

[0114] Define the preset range, which is recommended to be 30-50m (adjustable according to data density). The setting is based on the lateral migration distance of ore-forming elements in shallow-covered areas (≤30m) + sampling error (±5-10m).

[0115] Define the scope of aggregated data, including all geological structure data, topographic data, and environmental data (including dynamic environmental data) within a preset area surrounding the aggregation center.

[0116] Subsequently, the corresponding aggregation logic is executed according to the data type, strictly adhering to the preset value retrieval strategy:

[0117] Geochemical data are typically aggregated using their own sampling coordinates as the center. If there are duplicate geochemical sampling points within the aggregation range (such as multiple samplings from the same area), they are aggregated by type. For example, numerical values ​​(e.g., deep Au content: 12ppb, 14ppb, 13ppb) are aggregated to an arithmetic mean of 13ppb; if there are no duplicates, the original values ​​are retained.

[0118] Geological structural data is extracted, and vector data such as faults and rock masses within the aggregation range are extracted, quantified, and then aggregated. An example of the value selection strategy is given: numerical data (e.g., fault density: 0.6 faults / km). 2 0.8 lines / km 2 → Average 0.7 lines / km 2 ; By type (e.g., stratigraphic lithology: 3 granites, 1 sandstone) → granite, which has the highest proportion.

[0119] Topographic data is extracted from raster data such as DEMs, and topographic parameters within the aggregation range are extracted or integrated from borehole-exposed overburden data. For example, the value selection strategy is as follows: numerical data (e.g., overburden thickness: 15m, 20m, 18m) → average value 17.7m; categorized data (e.g., slope aspect: 4 sunny, 1 shady) → the most prevalent sunny slope (code 1).

[0120] Environmental data is integrated and aggregated from groundwater, vegetation, climate, and other data within the scope. Dynamic data is averaged over time before aggregation. For example, the value selection strategy is as follows: numerical data (e.g., groundwater pH values: 5.8, 6.2, 6.0) → average value 6.0; categorized data (e.g., vegetation types: 5 coniferous forests, 2 broadleaf forests) → coniferous forests (code 1) with the highest percentage.

[0121] Each aggregation center corresponds to one aggregation dataset, formatted as a CSV structured table, with each data entry containing:

[0122] Core identifiers: Aggregate center coordinates (consistent with the corresponding geochemical data sampling coordinates), aggregated dataset ID;

[0123] Multi-source aggregated fields: All quantified / coded indicators from geochemical data (after aggregation), geological structural data (after aggregation), topographic and geomorphological data (after aggregation), and environmental data (after aggregation).

[0124] For example, the aggregated dataset ID is J001, the aggregated center coordinates (B, L) are 35°12′30″, 110°23′15″, the geochemical data is Au=13ppb, and the geological structure data shows a fault density of 0.7 faults / km. 2 Topographic data: Overburden thickness = 17.7m; Environmental data: Groundwater pH = 6.0.

[0125] Aggregate dataset integrity verification and missing value imputation:

[0126] The integrity verification standard is as follows: the missing rate of core indicators (deep Au content, tectonic mineralization index, overburden thickness, and groundwater pH value) is ≤10%; the missing rate of non-core indicators (such as vegetation type and annual precipitation) is ≤30%; the verification tool is to use the isnull().sum() function of the Python Pandas library to count the missing fields of each data.

[0127] Differentiated missing value imputation rules (replacing uniform default values ​​to avoid distortion) set specific default values ​​according to data type and address meaning, rejecting meaningless imputation. When the data type is numeric (0 value with no geological significance), the missing value imputation rule uses the "mean value of the same geological unit" (such as the mean pH value of groundwater in the same fault zone, or the mean thickness of the overburden layer in the same stratum). When the data type is numeric (0 value with geological significance), the missing value imputation rule fills in 0 values ​​(representing "no such geological condition"). When the data type is categorical / coded, the missing value imputation rule fills in the "unknown category code" (such as 999) to avoid confusion with the mode.

[0128] Aggregate datasets and match them with the label space of mining sites:

[0129] Before matching, the mineral point labels are preprocessed: the coordinates of all mineral point labels are unified to the 2000 national geodetic coordinate system, consistent with the coordinate format of the aggregated dataset, and the attributes of each mineral point label are clearly defined (such as mineralization status: 1 = mineralized / 0 = non-mineralized, or mineralization potential of 4.5 points); the matching threshold is set: 50m is recommended (to adapt to the aggregation range) to ensure the spatial correlation between the aggregated dataset and the mineral point labels.

[0130] Specifically, for each aggregated dataset (coordinates P1 (B1, L1)) and all mining point labels (coordinates P2 (B2, L2)), the planar straight-line distance D is calculated using the geodetic coordinate distance formula:

[0131] D 1,2 =R×arccos[sinB1sinB2+cosB1cosB2cos(L1-L2)];

[0132] Where R is the Earth's radius, approximately 6371 km, and the result is converted to meters.

[0133] The tool uses the Python geopy library for fast calculations.

[0134] Matching rule execution: Prioritize matching, filter out all mining point labels that are D≤50m away from the aggregated dataset, and bind the nearest mining point label according to the "proximity principle"; each mining point label can be bound to multiple aggregated datasets, but each aggregated dataset is bound to only 1 mining point label, and there is no duplicate matching.

[0135] Generate an aggregated dataset - a mineral deposit label matching table (CSV format).

[0136] This embodiment achieves the integration and precise spatial correlation of geochemical data, geological structural data, topographic data, and environmental data through the aforementioned data aggregation and matching methods, significantly improving the accuracy and reliability of mineral deposit prediction. It not only solves the challenge of unified processing of multi-source heterogeneous data but also ensures that each aggregated dataset accurately reflects the comprehensive geological conditions of its corresponding region through meticulous data preprocessing, reasonable missing value imputation, and scientific spatial matching strategies. In particular, the method effectively improves the efficiency of mineral exploration and the accuracy of resource location by accurately calculating the distance between the aggregated dataset and the mineral deposit labels using geodetic coordinate distance formulas and achieving optimal matching based on the "proximity principle." Furthermore, this method emphasizes the understanding and respect for geological significance, avoiding biases caused by unreasonable data imputation, and providing solid data support and scientific basis for subsequent mineral resource assessment, metallogenic prediction, and exploration deployment. Finally, the generated structured matching table provides decision-makers with a clear and reliable reference framework, helping to formulate more accurate and effective exploration strategies.

[0137] Furthermore, in a specific embodiment, based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structural data, topographic data, and environmental data corresponding to each mineral deposit label, and the causal path of mineralization between these data and the mineral deposit label, including:

[0138] Based on the PCMCI algorithm, each geochemical data, geological structure data, topographic data and environmental data in each aggregated dataset is transformed into corresponding map nodes;

[0139] Based on each mineral point label, all map nodes in each aggregated dataset, and a pre-set set of geological prior constraints, the mineralization causal path between each aggregated dataset and the corresponding mineral point label is obtained.

[0140] Furthermore, based on each mineralization point label, all map nodes in each aggregated dataset, and a pre-set set of geological prior constraints, the mineralization causal path between each aggregated dataset and its corresponding mineralization point label is obtained, including:

[0141] Based on all the graph nodes in each aggregated dataset and a pre-set set of geological prior constraints, the causal value between every two graph nodes is obtained. Based on the causal value between every two graph nodes and a pre-set causal threshold, a corresponding graph edge is generated between two graph nodes whose causal value is greater than the causal threshold. Each graph edge has a corresponding causal direction and causal strength.

[0142] Based on the causal direction corresponding to each graph edge in each aggregated dataset, all the top-level and bottom-level graph nodes in each aggregated dataset are selected, and graph edges are formed between each top-level graph node and the corresponding mineral point label in the aggregated dataset; among them, the connection relationship between the bottom-level graph node and the corresponding mineral point label in each aggregated dataset is the mineralization causal path.

[0143] Specifically, graph node transformation (aggregated data → nodes):

[0144] Each aggregated dataset directly maps multi-source data to nodes in a historical mineralization causal graph, ensuring a one-to-one correspondence between nodes and data with clearly defined attributes.

[0145] First, the transformation rules for nodes are based on the "quantitative / encoded indicators" in the aggregated dataset. Each indicator corresponds to one map node, and the node attributes include "data type, variable name, aggregated value, and sampling coordinates." Node naming conventions: use the format "data type-variable name" (e.g., "Geochemistry-deep Au content," "Geological structure-fault density") to avoid duplicate names. Special handling: categorized / encoded data (e.g., slope aspect = 1 (sunward)) needs to have its encoding meaning annotated in the node attributes to ensure the algorithm can recognize it during causal calculations. Example:

[0146] Geochemistry - Deep Au Content (13ppb), Spectrum Node Name: Geochemistry - Deep Au Content, Node Attributes: Data Type, Geochemistry, Variable Value, 13ppb, Unit, ppb, Coordinates: 35°12′30″, 110°23′15″.

[0147] Graph edge generation (node ​​→ edge with attributes):

[0148] Based on the PCMCI algorithm and geological prior constraint sets, the causal relationships between nodes are calculated, graph edges are generated, and the causal direction and intensity are determined.

[0149] Core input pre-parameter settings: input data, all map nodes (including variable values) of a single aggregate dataset, and a pre-set set of geological prior constraints (causal direction constraints, conditional independence constraints, and variable priority constraints); key parameters, causal threshold (recommended 0.3-0.4, which can be adjusted through the validation set), used to filter valid causal relationships; PCMCI algorithm parameters (significance level α=0.05, maximum lag order k=2).

[0150] Causality calculation and constraint application:

[0151] Node causality is calculated only: the conditional independence test of the PCMCI algorithm is performed, and the significance p-value of the association is calculated for all node pairs (such as geochemistry-deep Au content → environment-groundwater pH value), and node pairs with p < 0.05 are retained; causality value is calculated, and for significantly associated node pairs, the average treatment effect (ATE) is calculated through the "Causal Effect Estimation" module of PCMCI, and the ATE value is normalized to the [0,1] interval as the causality value between nodes (the closer the causality value is to 1, the stronger the causal association).

[0152] Geological prior constraints are used for filtering, causal direction constraints are used to correct the causal flow direction of node pairs (e.g., by constraining "deep Au content → surface As anomaly", the algorithm's initial judgment of "surface As anomaly → deep Au content" is corrected to the correct direction); conditional independence constraints are used to disconnect the association when the constraint condition is triggered (e.g., when "overburden thickness ≥ 80m", the association of "geochemistry - deep Au content → geochemistry - surface Au content" is disconnected, and the causal value is set to 0); variable priority constraints are used to discount the causal values ​​of low-priority nodes (e.g., "topography - slope" with a weight < 0.2) with other nodes according to their weights (e.g., the original causal value 0.4 × weight 0.1 = 0.04) to reduce interference.

[0153] Graph generation: Filter valid edges, retain node pairs with causal values ​​> causal threshold (e.g., 0.3), and generate corresponding graph edges; assign edge attributes, binding two core attributes to each edge: causal direction (according to the constraint-corrected "cause node → result node") and causal strength (i.e., the normalized causal value); handle invalid edges, node pairs with causal values ​​≤ the threshold or those that are broken by constraints will not generate graph edges.

[0154] Node-level filtering and mineral point tag association (generating ore-forming causal paths):

[0155] By filtering node levels according to causal direction, the top-level node is associated with the mineralization point label to clarify the causal path of mineralization.

[0156] Top-level / bottom-level graph node selection: Bottom-level graph nodes are nodes in the current aggregated dataset whose nodes have no other nodes pointing to them (i.e., no incoming edges, only outgoing edges), representing the "initial driving factor" of mineralization; top-level graph nodes are nodes in the current aggregated dataset whose nodes have no other nodes pointing to them (i.e., no outgoing edges, only incoming edges), representing the "direct indicator factor" of mineralization.

[0157] The mining point label corresponding to each aggregated dataset is used as the "final result node" to generate new graph edges with all the top-level nodes; edge attributes are assigned, and the causal direction of the new edge is fixed as "bottom-level node → mining point label". The causal strength is calculated according to the total strength of the causal path from the bottom-level node to the top-level node (total strength = weighted product of the strengths of each edge in the path); in special cases, if the aggregated dataset is bound to "non-mineralized label (0)", the causal strength of the new edge is set to 0, which means there is no effective mineralized path.

[0158] Mineralization causal path extraction: The path definition is a complete link connecting the lowest-level node and the mineralization point label. Path filtering retains paths with total strength greater than the causal threshold and removes low-intensity invalid links.

[0159] This embodiment accurately transforms aggregated multi-source geological data into map nodes and, guided by geological prior constraints, rigorously calculates the causal direction and intensity between nodes using the PCMCI algorithm, effectively constructing mineralization causal paths with clear geological significance. It not only overcomes the shortcomings of traditional data-driven methods, such as susceptibility to spurious correlations and lack of mechanistic support, but also clearly reconstructs the complete mineralization logic chain from deep material sources to surface anomaly responses through hierarchical modeling of "lowest level—highest level—mineral point label." Simultaneously, geological prior constraints (such as causal direction correction, conditional independence disconnection, and variable priority weighting) significantly improve the scientific rigor and reliability of the causal map, preventing the algorithm from drawing erroneous inferences that violate geological laws. The final extracted high-intensity mineralization causal paths provide highly discriminative causal guidance features for target area prediction models and a visualized and interpretable knowledge base for geologists to understand mineralization control mechanisms and optimize exploration deployment, greatly enhancing the accuracy and practicality of predicting concealed gold deposits in shallowly covered areas.

[0160] Optionally, in a specific embodiment, such as Figure 3 As shown, the geological prior constraint set is obtained through the following steps:

[0161] S1. Extract objects from the pre-set geological mineralization expert theory to obtain multiple variable objects, and convert all variable objects into recognizable objects of the PCMCI algorithm;

[0162] S2. Based on pre-set constraint rules, apply rule constraints to all identifiable objects to obtain a set of geological prior constraints; among which, rule constraints include: causal direction constraints, conditional independence constraints, and variable priority constraints.

[0163] Furthermore, the variable objects include numerical geological variable objects or categorical / qualitative geological variable objects;

[0164] Then, all variable objects are converted into objects recognizable by the PCMCI algorithm, including:

[0165] Based on each numerical geological variable object and the pre-set quantitative indicators for each numerical geological variable object, each numerical geological variable object is quantitatively processed.

[0166] Based on each category / qualitative geological variable object and a pre-set coding strategy, obtain the coded value corresponding to each category / qualitative geological variable object;

[0167] The object format is unified for each quantified numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object to obtain the recognizable object of the PCMCI algorithm for each variable object;

[0168] The object format is unified as follows: each quantified numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object are stored in a structured form to adapt to the input format of the PCMCI algorithm.

[0169] Specifically, the first step is to extract the variable objects from the geological mineralization expert theory. This involves breaking down the "numerical geological variable objects" and "classification / qualitative geological variable objects" directly related to mineralization from the mineralization-specific expert theory for gold deposits in shallow-covered areas, ensuring that the variable objects cover the entire chain.

[0170] The extraction basis focuses on theories of ore-forming material sources (such as granite body ore supply and deep ore source layer), migration channel theories (such as fault-guided ore and rock mass contact zone ore control), element migration theories (such as nano gold migration and vertical migration of active elements), and environmental regulation theories (such as groundwater pH regulation and the influence of overburden thickness).

[0171] Variable object extraction results (categorized presentation): Numerical geological variable objects (deep Au content, surface As anomaly intensity, fault density, tectonic ore-conducting capacity index, overburden thickness, groundwater pH value, nano-gold migration index, rock mass contact zone distance, etc.), based on the theories of ore-forming material source, ore-conducting channels, element migration, and environmental regulation; Categorical / qualitative geological variable objects (strata lithology, slope aspect, vegetation type, geomorphic unit, fault properties (ore-conducting / non-ore-conducting), rock mass type (granite / volcanic rock)), based on the theories of stratigraphic control of ore, topographic influence on element migration, and rock mass mineralization specificity.

[0172] Secondly, the variable objects are transformed into objects that the PCMCI algorithm can recognize. That is, the abstract variable objects are transformed into concrete objects that the algorithm can process in the order of numerical quantization → classification / qualitative encoding → format unification, so as to avoid constraint failure due to format incompatibility.

[0173] The quantification of numerical geological variables involves setting clear "quantification indicators, units, and value ranges" for extracted numerical variables based on their metallogenic geological significance, ensuring the scientific traceability of the quantification results. For example, for deep Au content, the pre-set quantification indicator is the test value of borehole core samples (ICP-MS method), and the quantification operation is to directly use the test data and retain one decimal place; if multiple test values ​​exist (multiple core samples from the same sampling point), the arithmetic mean is taken. For surface As anomaly intensity, the pre-set quantification indicator is the As content in soil / river sediments ÷ regional background value, and the quantification operation is to calculate the anomaly multiple and retain two decimal places.

[0174] The coding of categorical / qualitative geological variables involves assigning unique codes based on the geological significance of the variables. This ensures that the codes are not duplicated and reflect the mineralization differences between categories. The coding strategy prioritizes "ordered coding" (aligned with mineralization contribution). For example, for stratigraphic lithology, the pre-set quantitative index is ordered coding according to the degree of mineralization favorability: granite (most favorable) = 1, sandstone (relatively favorable) = 2, shale (average) = 3, mudstone (unfavorable) = 4. The quantitative processing operation matches the corresponding code value based on the stratigraphic lithology type marked on the geological map. For slope aspect, the pre-set coding strategy is coded according to the favorability of element migration: sunny slope (favorable for evaporation anomalies) = 1, shady slope = 2, flat land = 3. The coding processing operation extracts the slope aspect based on the DEM and classifies it according to azimuth (0°-180° is sunny slope, 180°-360° is shady slope, and slope < 3° is flat land).

[0175] Then, the quantified numerical variable objects and the encoded classification / qualitative variable objects are stored in a unified "structured form" to ensure that each identifiable object contains "core attributes + geological meaning" and is compatible with the input format of the PCMCI algorithm.

[0176] Subsequently, constraints are applied to all algorithm-identifiable objects according to three categories of rules: "causal direction constraints, conditional independence constraints, and variable priority constraints," ultimately forming a structured set of geological prior constraints to ensure that the constraints can be directly invoked by the PCMCI algorithm.

[0177] Causal direction constraints (clarifying the flow of "cause node → result node") are based on the logic of "ore-forming factors → intermediate process → ore-forming result", setting the causal flow between identifiable objects and correcting possible causal inversions in the algorithm.

[0178] Conditional independent constraints (setting a prerequisite threshold for the constraint to take effect) are based on the "effective range" of mineralization conditions, setting threshold conditions for identifiable objects, and disconnecting associations that do not conform to geological laws.

[0179] Variable priority constraints (quantifying the contribution of variables to mineralization): weight coefficients (0-1 range) are set according to the degree of correlation between identifiable objects and mineralization. The algorithm prioritizes retaining the causal relationships of high-weight variables.

[0180] Finally, the three types of constraints are integrated in JSON format to form a structured set of geological prior constraints that can be directly input into the PCMCI algorithm.

[0181] This embodiment extracts variable objects from geological and metallogenic expert theories and transforms these variable objects into a form recognizable by the PCMCI algorithm. Simultaneously, it applies pre-set constraint rules to obtain a structured set of geological prior constraints. This not only ensures that data processing and analysis strictly adhere to geological principles and knowledge, improving the scientific rigor and reliability of causal relationship inferences, but also effectively overcomes the spurious correlation problem present in traditional data analysis methods, enhancing the accuracy and practicality of the metallogenic prediction model. Furthermore, by quantifying and encoding numerical and categorical / qualitative geological variable objects separately, and then unifying the format to adapt to the PCMCI algorithm input, efficient integration and utilization of multi-source heterogeneous geological data are achieved. The resulting set of geological prior constraints significantly improves the performance of the PCMCI algorithm in exploring causal relationships between complex factors such as geochemistry and geological structure, providing strong support for the prediction of concealed ore bodies in shallowly covered areas. This enables geologists to more accurately identify potential mineral deposits, optimize exploration strategies, reduce exploration costs, and increase the success rate of mineral exploration.

[0182] Optionally, in a specific embodiment, causal feature extraction is performed on the historical mineralization causal map to obtain corresponding causal guidance features, including:

[0183] Feature extraction, feature integration, and standardization are performed on all causal paths corresponding to each mineralization point label in the historical metallogenic causal map to obtain causal guidance features corresponding to each mineralization point label in the historical metallogenic causal map. Causal guidance features include: intensity weighted features, causal interaction features, causal contribution features, and causal existence features.

[0184] Specifically, firstly, weighted features (i.e., intensity-weighted features) of core causal paths are extracted based on the "core causal paths" (such as "deep Au content → structural ore-guiding capacity index → ​​ore point label") in historical metallogenic causal maps. By weighting the path intensity, the metallogenic indicative role of core variables is strengthened. From the historical metallogenic causal maps, core causal paths with a total intensity ≥ 0.4 (adjustable) are retained, while low-intensity spurious paths (such as "slope → precipitation → ore point label" with a total intensity < 0.3) are removed. For each variable, the intensity-weighted sum of all core paths it participates in is calculated (the weight is the total intensity of the path), using the formula: variable weighted value = ∑ (total intensity of core paths in which the variable participates × segment intensity of the variable in the path). Feature standardization is then performed by Z-score standardization of the weighted values ​​of all variables to generate intensity-weighted features.

[0185] Secondly, causal interaction features (reflecting the synergistic mineralization effect of multi-source data) are constructed based on the "multivariate transmission relationship" (such as "geological structure variable → geochemical variable → mineral occurrence label") in historical mineralization causal maps to capture the synergistic effect of multiple factors in the mineralization process. From the core causal path, continuously transmitted variable pairs (such as "tectonic mineralization capacity index → ​​nano-gold migration index", "groundwater pH value → surface As anomaly intensity") are extracted, prioritizing variable pairs of different data types (such as geological structure + geochemistry, dynamic environment + geochemistry). For each pair of synergistic variables, "multiplicative interaction" or "additive interaction" is used to construct features, with multiplicative interaction preferred. The correlation between the interaction features and mineral occurrence labels is calculated, retaining features with a correlation coefficient ≥ 0.5 and eliminating redundant interaction terms.

[0186] The causal contribution feature (quantifying the mineralization importance of a single variable) is based on the edge attributes (causal strength weight) of the historical mineralization causal graph. It quantifies the direct and indirect causal contributions of each variable to the mineralization label, clarifying the mineralization priority of the variable. If there is a direct causal edge between the variable and the mineralization label (e.g., nano-gold migration index → ​​mineralization label), then the direct causal contribution = the strength weight of the edge; if there is no direct causal edge, then it is 0. If the variable indirectly points to the mineralization label through an intermediate variable (e.g., tectonic ore-conducting capacity index → ​​nano-gold migration index → ​​mineralization label), then the indirect causal contribution = the product of the strengths of each segment in the path (e.g., 0.90 × 0.78 = 0.702). The total contribution = direct causal contribution + indirect causal contribution, normalized to the [0,1] interval, is used as the causal contribution feature.

[0187] The existence feature of causal paths (marking key mineralization conditions) is constructed based on whether variables participate in core paths in historical mineralization causal maps. This constructs binary marking features to reflect whether a variable is a necessary condition for mineralization, assisting the model in identifying favorable combinations of mineralization conditions. If a variable participates in ≥2 core causal paths (total intensity ≥0.4), it is marked as a "key mineralization variable" with an eigenvalue of 1; otherwise, it is marked as a "non-key variable" with an eigenvalue of 0. For independent constraints (such as "overburden thickness < 80m" or "structural ore-conducting capacity index ≥0.8"), a condition satisfaction feature is constructed. If a variable satisfies the constraint (such as overburden thickness = 50m < 80m), the eigenvalue is 1; otherwise, it is 0. The key variable markings are combined with path condition markings to generate "mineralization condition combination features" (such as "key variable exists + constraint satisfied = 1 + 1 = 2" or "key variable exists + constraint not satisfied = 1 + 0 = 1").

[0188] This embodiment systematically extracts causal features from historical mineralization causal maps, constructing causal-guided features covering four dimensions: intensity weighting, causal interaction, causal contribution, and causal existence. This significantly improves the discriminative ability and geological interpretability of the gold ore target area prediction model.

[0189] In addition, the target area prediction model in this application embodiment can be selected from: Random Forest (RF), Gradient Boosting Decision Tree (XGBoost / LightGBM / CatBoost), Logistic Regression (LR) + Feature Engineering Enhancement Framework, Support Vector Machine (SVM) or Graph Neural Network (GNN, such as GCN / GAT).

[0190] Furthermore, this application employs a causal inference enhancement model (PCMCI+XGBoost / SHAP-LightGBM), using a core architecture of causal mining → feature selection → predictive modeling → interpretability verification → constraint feedback. It deeply integrates the causal logic of PCMCI, the predictive power of gradient boosting trees, and the interpretability of SHAP, ensuring the model "understands geological causality and can accurately predict," while embedding geological prior constraints throughout the process to avoid learning spurious associations. Specifically, this includes: PCMCI causal mining and constraint embedding (causal logic layer), causal feature selection and enhancement (core feature engineering layer), XGBoost / LightGBM predictive modeling (accurate prediction layer), SHAP interpretability verification and constraint validation (logic guarantee layer), and model optimization and target area output (closed-loop optimization and application layer). PCMCI causal mining and constraint embedding (causal logic foundation layer) uses historical mineralization data and geological prior constraint sets to mine real mineralization causal paths using the PCMCI algorithm, outputting structured causal knowledge (causal relationships, causal strength, effective paths), providing "causal priors" for subsequent modeling. The causal feature selection and enhancement (core feature engineering layer) uses causal knowledge output by PCMCI to select features with "real causal relationships," eliminates irrelevant features, and constructs causal-oriented enhancement features to provide a "high-quality causal feature set" for predictive modeling. The XGBoost / LightGBM predictive modeling (accurate prediction layer) takes the "final causal feature set" as input and uses a gradient boosting tree model to learn the nonlinear mapping relationship between causal features and mineral deposit labels, outputting target area prediction results (metallogenic probability / metallogenic potential score), while embedding geological prior constraints to optimize model training. The SHAP interpretability verification and constraint validation (logic assurance layer) uses the SHAP tool to quantify the contribution of each causal feature to the prediction result, verifying whether the model's prediction logic is consistent with the causal relationship of PCMCI and geological prior constraints. If conflicts exist, feedback correction is triggered. The model optimization and target area output (closed-loop optimization and application layer) iteratively optimizes model parameters / feature sets based on the validation results, ultimately outputting accurate and interpretable target area prediction results, and generating a report that can be reviewed by geological experts.

[0191] In addition, this application provides a gold ore target area prediction system based on mineralization causality, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the above-mentioned gold ore target area prediction method based on mineralization causality.

[0192] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0193] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0194] In this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first and second features are in direct contact, or that they are in indirect contact through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0195] In the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0196] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make modifications, alterations, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A gold deposit target area prediction method based on mineralization causality guidance, characterized in that, The method is applied to gold target area prediction in shallowly covered areas, and the method includes: Obtain geophysical data of the region to be predicted, and input the geophysical data into a pre-trained target area prediction model to obtain the prediction result of whether the target area is the target area; The target area prediction model is trained through the following steps: Acquire historical data for the target area, including historical geophysical data and multiple known mineral deposits within the target area, with each known mineral deposit having a corresponding known mineral deposit label; Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data and obtain the corresponding historical mineralization causal map. Causal features were extracted from historical mineralization causal maps to obtain corresponding causal guidance features; The causal guidance features are input into the target area prediction model, and the model hyperparameters are optimized by the gradient descent algorithm to obtain a target area prediction model for gold mine target area prediction in shallow-covered areas.

2. The gold deposit target area prediction method based on mineralization causality guidance according to claim 1, characterized in that, The historical geophysical data includes all geochemical data, geological structure data, topographic data, and environmental data within the target area; Based on a pre-defined set of geological prior constraints, the PCMCI algorithm is used to mine causal relationships in historical data, obtaining corresponding historical mineralization causal maps, including: Match all geochemical data, geological structure data, topographic data, and environmental data within the target area with all mineral deposit labels; Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structure data, topographic data and environmental data corresponding to each mineral deposit label, and the causal path of mineralization between the mineral deposit label and the label. Based on the mineralization causal path corresponding to each mineralization point label, a corresponding historical mineralization causal map is obtained.

3. The gold deposit target area prediction method based on mineralization causality guidance according to claim 2, characterized in that, Each geochemical data, geological structure data, topographic data, and environmental data has corresponding sampling coordinates; each mineral deposit label has corresponding mineral deposit coordinates. The geochemical data, geological structure data, topographic data, and environmental data within the target area are matched with all mineral deposit labels, including: Based on the sampling coordinates corresponding to each geochemical data, geological structure data, topographic data, and environmental data within the target area, and using the sampling coordinates of each geochemical data as the aggregation center, all geological structure data, topographic data, and environmental data within a preset range are aggregated to obtain multiple aggregated datasets within the target area. During the aggregation process, when multiple data of the same type appear, sampling is performed based on a predefined value selection strategy. The value selection strategy is as follows: when the data type of the same type of data is numerical data, take the arithmetic mean of all data in the same type of data; when the data type of the same type of data is categorical / coded data, take the data with the highest proportion in the same type of data. Verify the integrity of each aggregated dataset and fill in missing data based on pre-set default values; Based on the fact that each aggregated dataset contains geochemical data with corresponding sampling coordinates and each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels.

4. The gold deposit target area prediction method based on mineralization causality guidance according to claim 3, characterized in that, Based on the fact that each aggregated dataset contains geochemical data with corresponding sampling coordinates and each mineral point label has corresponding mineral point coordinates, all aggregated datasets within the target area are matched with all mineral point labels, including: Based on the sampling coordinates corresponding to the geochemical data in each aggregated dataset and the corresponding mineral point coordinates for each mineral point label, and using a pre-set formula (Formula 1), the distance between each aggregated dataset and each mineral point label is obtained; Formula 1 is: D n,m =R×arccos[sinB n sinB m +cosB n cosB m cos(L n -THE m )]; Where n is the index of the aggregated dataset, m is the index of the mining point label, (B n ,L n ) represents the sampling coordinates of the aggregated dataset with index n, (B m ,L m ) represents the coordinates of the mineral point labeled with index m, R is the Earth's radius, and D is the Earth's coordinates. n,m The distance between the aggregate dataset with index n and the mining point labels with index m; Based on the distance between each aggregated dataset and each mining point label and the pre-set proximity principle, all aggregated datasets within the target area are matched with all mining point labels.

5. The gold deposit target area prediction method based on mineralization causality guidance according to claim 3, characterized in that, Based on a pre-set set of geological prior constraints, the PCMCI algorithm is used to obtain all geochemical data, geological structural data, topographic data, and environmental data corresponding to each mineral deposit label, and the causal path of mineralization between these data and the mineral deposit label, including: Based on the PCMCI algorithm, each geochemical data, geological structure data, topographic data and environmental data in each aggregated dataset is transformed into corresponding map nodes; Based on each mineral point label, all map nodes in each aggregated dataset, and a pre-set set of geological prior constraints, the mineralization causal path between each aggregated dataset and the corresponding mineral point label is obtained.

6. The gold deposit target area prediction method based on mineralization causality guidance according to claim 5, characterized in that, Based on each mineral deposit label, all map nodes in each aggregated dataset, and a pre-defined set of geological prior constraints, the mineralization causal path between each aggregated dataset and its corresponding mineral deposit label is obtained, including: Based on all the graph nodes in each aggregated dataset and a pre-set set of geological prior constraints, the causal value between every two graph nodes is obtained. Based on the causal value between every two graph nodes and a pre-set causal threshold, a corresponding graph edge is generated between two graph nodes whose causal value is greater than the causal threshold. Each graph edge has a corresponding causal direction and causal strength. Based on the causal direction corresponding to each graph edge in each aggregated dataset, all the top-level and bottom-level graph nodes in each aggregated dataset are selected, and graph edges are formed between each top-level graph node and the corresponding mineral point label in the aggregated dataset; among them, the connection relationship between the bottom-level graph node and the corresponding mineral point label in each aggregated dataset is the mineralization causal path.

7. The gold deposit target area prediction method based on mineralization causality as described in claim 1, characterized in that, The set of geological prior constraints is obtained through the following steps: Object extraction is performed on the pre-set geological mineralization expert theory to obtain multiple variable objects, and all variable objects are transformed into recognizable objects by the PCMCI algorithm; Based on pre-set constraint rules, rule constraints are applied to all identifiable objects to obtain a set of geological prior constraints; wherein, the rule constraints include: causal direction constraints, conditional independence constraints, and variable priority constraints.

8. The gold deposit target area prediction method based on mineralization causality guidance according to claim 7, characterized in that, The variable objects include numerical geological variable objects or categorical / qualitative geological variable objects; Convert all variable objects into objects recognizable by the PCMCI algorithm, including: Based on each numerical geological variable object and the pre-set quantitative indicators for each numerical geological variable object, each numerical geological variable object is quantitatively processed. Based on each category / qualitative geological variable object and a pre-set coding strategy, obtain the coded value corresponding to each category / qualitative geological variable object; The object format is unified for each quantified numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object to obtain the recognizable object of the PCMCI algorithm for each variable object; The unified object format is as follows: each quantized numerical geological variable object, each categorical / qualitative geological variable object, and the corresponding encoded value of each categorical / qualitative geological variable object are stored in a structured form to adapt to the input format of the PCMCI algorithm.

9. The gold deposit target area prediction method based on mineralization causality as described in claim 1, characterized in that, The historical metallogenic causal map includes multiple metallogenic causal pathways; Causal features were extracted from historical mineralization causal maps to obtain corresponding causal guidance features, including: Feature extraction, feature integration, and standardization are performed on all causal paths corresponding to each mineralization point label in the historical mineralization causal map to obtain causal guidance features corresponding to each mineralization point label in the historical mineralization causal map; the causal guidance features include: intensity weighted features, causal interaction features, causal contribution features, and causal existence features.

10. A gold ore target area prediction system based on mineralization causality guidance, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the gold ore target area prediction method based on mineralization causality as described in any one of claims 1 to 9.