A deep geothermal resource prediction method based on mechanism-space-distribution fusion
Patent Information
- Application Number
- CN202610430777.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-08-18
AI Technical Summary
[0007]本发明的目的在于克服现有深部地热资源预测方法中物理机理缺失、空间信息忽视、标签依赖强、漏报率高的技术缺陷,提供一种基于“机理-空间-分布”多维异构融合架构(Mechanism-Space-Distribution Heterogeneous Fusion Architecture,MSD-Fusion)的深部地热资源预测方法
(1)可解释性与预测精度协同提升:通过机理驱动层注入地质先验,构建“机理-空间-分布”融合体系,破解传统机器学习“黑箱”难题,多维特征互证增强地质合理性,在蓝田地热勘查中实现AUC=0.98、召回率=1.00的优异性能,达成零漏报与高精度的双重目标。
Smart Images

Figure CN122595170A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological big data mining and artificial intelligence, and in particular to a method for predicting deep geothermal resources based on mechanism-space-distribution fusion. Background Technology
[0002] Deep geothermal resources, as a clean, stable, and sustainable new energy source, are of significant strategic importance for promoting the transformation of the national energy structure and achieving the "dual-carbon" goal through precise exploration and efficient development. The formation of geothermal resources is a complex process involving the coupling of multiple physical fields, including heat, fluidity, mechanics, and chemistry. Surface geochemical anomalies, as direct responses of deep thermal-mass migration activities at the shallow surface, are key clues for locating concealed geothermal fields. Geothermal resource prediction based on geochemical data has become a research hotspot in the industry.
[0003] With the development of the data-driven scientific paradigm, existing geothermal resource prediction technologies are mainly divided into two categories: traditional geostatistical methods and conventional machine learning methods. Although some progress has been made, many technical bottlenecks still exist, making it difficult to meet the exploration needs of deep geothermal resources with high accuracy and low underreporting.
[0004] Traditional geostatistical methods, represented by single-element anomaly delineation, correlation analysis, and principal component analysis (PCA), rely heavily on geological experts' experience to set thresholds (such as μ+2σ) for anomaly identification. These methods have significant limitations: ① They are constrained by linear assumptions, making it difficult to capture nonlinear geochemical distribution patterns caused by complex deep hydrothermal activity, and they are poorly adaptable to complex geological conditions; ② They ignore the spatial continuity and topological structure of the geochemical field, treating sampling points as independent and identically distributed samples, thus failing to accurately reflect the spatial correlation of hydrothermal migration; ③ Threshold setting is highly subjective, easily influenced by differences in expert experience, and lacks stability and generalization ability.
[0005] Conventional machine learning methods, such as Support Vector Machines (SVM) and Artificial Neural Networks (ANN), while superior to traditional methods in nonlinear fitting capabilities, still face severe challenges in geothermal exploration scenarios: ① The lack of physical mechanisms leads to a "black box" problem. Purely data-driven models lack constraints from geological mineralization mechanisms, are prone to learning spurious correlations in the data, have poor model interpretability, and are difficult to gain the approval of geological experts; ② The ambiguous definition of negative samples causes label noise problems. Geothermal exploration is a typical scenario of "positive samples are known, negative samples are potential." Traditional supervised learning forcibly labels undiscovered areas as "non-mineralized," leading to a shift in the decision boundary and increasing the risk of underreporting; ③ The dilemma of high-dimensional small samples is prominent. Geological exploration samples are costly and time-consuming to obtain, often facing an imbalance between hundreds of elemental features and only a few hundred samples, which easily leads to model overfitting and limited generalization ability.
[0006] To address the shortcomings of existing technologies, there is an urgent need to construct a multi-dimensional heterogeneous fusion prediction framework that integrates geological priors, spatial characteristics, and distribution patterns. This framework would overcome the limitations of single-perspective models and enable robust prediction and zero-miss screening of deep geothermal resources in complex geological contexts characterized by small samples, high noise, and nonlinearity, thus providing precise technical support for geothermal exploration projects. Summary of the Invention
[0007] The purpose of this invention is to overcome the technical shortcomings of existing deep geothermal resource prediction methods, such as lack of physical mechanisms, neglect of spatial information, strong label dependence, and high false negative rates. It provides a deep geothermal resource prediction method based on a Mechanism-Space-Distribution Heterogeneous Fusion Architecture (MSD-Fusion). This method innovatively constructs a multi-dimensional heterogeneous fusion feature system, combining dynamic feature denoising, ternary consensus integration, and asymmetric decision-making strategies to achieve high-precision, low-false-negative prediction of geothermal resources, while also possessing good geological interpretability.
[0008] This invention provides a method for predicting deep geothermal resources based on mechanism-space-distribution fusion, comprising: S1. Obtain geochemical multi-element analysis data and basic geological data of the target area, perform limit correction and standardization preprocessing on the geochemical multi-element analysis data, construct label vectors based on known geothermal anomaly information in the basic geological data, and divide the training set and test set. S2. Construct a multi-dimensional heterogeneous fusion feature system that coordinates the mechanism-driven layer, spatial perception layer and data distribution layer. Using standardized geochemical multi-element analysis data as input, generate multi-dimensional heterogeneous fusion features that integrate geothermal mineralization mechanism, geochemical field spatial structure and data distribution law. S3. A recursive feature elimination algorithm is adopted, with random forest as the base learner and Gini importance as the evaluation index. Low-importance features are eliminated in each round, and the core geothermal geochemical fingerprint set is obtained through iterative screening. S4. A three-element consensus integrated prediction module is constructed based on gradient boosting tree, random forest and extreme random tree. The core geothermal geochemical fingerprint set is used as input and the label vector is used as supervision. An equal-weight soft voting strategy is adopted to fuse the prediction results of each model and output the joint prediction probability. S5. Based on the recall rate anchoring strategy, determine the asymmetric decision boundary. Based on the joint prediction probability and label vector on the training set, set the minimum recall rate constraint, optimize the optimal threshold, perform full-domain grid prediction on the target area, and output the geothermal resource potential prediction map and high-potential target area.
[0009] Furthermore, in step S2, the mechanism-driven layer takes standardized elemental concentration data as input and outputs radioactive heat-generating potential energy. Hydrothermal fluid tracer index and the thermal-mass migration coupling factor generated by the nonlinear interaction between the two. This constitutes a 3D mechanism feature vector; wherein, the radioactive heat generation potential energy Based on the decay heat generation characteristics of uranium, thorium, and potassium, the formula is: In the formula, , , These represent the elemental weights of uranium, thorium, and potassium, respectively. , , These represent the concentrations of uranium, thorium, and potassium, respectively; the hydrothermal fluid tracer index... The formula for polymerizing seven fluid indicator elements—arsenic, antimony, lithium, rubidium, cesium, tungsten, and tin—is as follows: The heat-mass migration coupling factor The product of the heat generation potential energy and the fluid tracer index is given by the formula: .
[0010] Furthermore, in step S2, the spatial perception layer takes the spatial coordinates of the sampling points and elemental concentration data as input, constructs a spatial topology based on the K-nearest neighbor graph, and outputs the neighborhood mean and variance at two scales, K=5 and K=15, as well as the fractal singularity index used to identify weak anomalous signals. This forms a 5-dimensional spatial feature vector; the formula for calculating the fractal singularity index is: In the formula, Indicates the element concentration at the sampling point. This represents the neighborhood mean at the corresponding scale. This indicates the smoothing term.
[0011] Furthermore, in step S2, the data distribution layer takes standardized element concentration data as input and uses the isolated forest algorithm for unsupervised training, outputting distribution anomaly scores to form a 1-dimensional data distribution feature vector. The isolated forest algorithm is set to 100 trees, 25 sample splits, and "auto" contamination level. The algorithm output decision function value is inverted so that a higher score represents a stronger degree of anomaly, serving as a complement to the prior probability and supervised learning features.
[0012] Furthermore, in step S3, the input of the recursive feature elimination algorithm is a multidimensional initial feature set composed of multidimensional heterogeneous fusion features output from the mechanism-driven layer, spatial perception layer, and data distribution layer, and multidimensional original element features. After each round of training, the Gini importance is calculated and the features ranked in the bottom 5% are removed. The feature set is iteratively updated until a 50-dimensional core geothermal geochemical fingerprint set is retained.
[0013] Furthermore, in step S4, the parameters of the gradient boosting tree model are configured as follows: learning rate = 0.05, number of trees = 100, maximum depth = 3, minimum number of leaf node samples = 3, subsample ratio = 0.8; the parameters of the random forest model are configured as follows: number of trees = 150, maximum depth = 5, minimum number of leaf node samples = 2, and class weight balance; the parameters of the extreme random tree model are configured as follows: number of trees = 150, maximum depth = 10, minimum number of leaf node samples = 2, random feature splitting, and class weight balance; the random seed for all models is set to 42.
[0014] Furthermore, in step S5, the recall anchoring strategy takes the joint predicted probability and label vector on the training set as input and sets a minimum recall constraint. Or adjust to meet the requirement of zero underreporting in geothermal exploration. The candidate thresholds are traversed in the range of 0.1 to 0.9 with a step size of 0.01 through grid search, and the optimal threshold that satisfies the recall constraint is output. If the preset recall constraint cannot be satisfied after traversal, the threshold is automatically switched to the F2 score maximization strategy to determine the threshold, where the adjustment factor β=2 in the F2 score.
[0015] Furthermore, in step S5, the global prediction takes the standardized feature data of all grid nodes in the study area as input, and obtains the probability field after processing in steps S2-S4. The optimal threshold is used for binarization judgment, and a geothermal resource potential prediction map and high potential target area with a resolution of 50m×50m are output. The contrast of the visualization color mark is optimized by histogram equalization, and the prediction results are superimposed and compared with the known fault tectonic zones in the target area for verification.
[0016] Furthermore, the method is applicable to scenarios such as deep geothermal resource exploration, dry hot rock target area selection, hydrothermal geothermal field boundary delineation, or pre-exploitation assessment of geothermal resources. It enables the prediction of deep, hidden geothermal resources based on shallow surface soil or rock geochemical data. The method also includes a model deployment process, which uses Joblib to save the complete artifact consisting of the trained ensemble model, normalizer, feature selector, and optimal threshold, supporting direct loading and deployment.
[0017] Furthermore, the method also includes a full-process visualization module, which automatically generates an architecture flowchart, a spatial singularity and RFE curve comparison chart, a recall rate anchoring decision boundary diagram, and a geothermal resource potential prediction map for technical process display and prediction result output.
[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) Synergistic improvement of interpretability and prediction accuracy: By injecting geological priors into the mechanism-driven layer, a fusion system of "mechanism-space-distribution" is constructed to solve the "black box" problem of traditional machine learning. Multi-dimensional feature mutual verification enhances geological rationality. In the geothermal exploration of Lantian, excellent performance of AUC=0.98 and recall=1.00 is achieved, and the dual goals of zero false negatives and high accuracy are achieved.
[0019] (2) Strong adaptability to complex geological scenarios: The multidimensional heterogeneous fusion feature system simultaneously captures the mechanism, space and distribution patterns. The RFE algorithm and the ternary integration mechanism enhance the model's noise resistance and generalization ability, which can effectively address core technical challenges such as high-dimensional small samples, label noise, and spatial heterogeneity.
[0020] (3) Optimization of exploration costs and benefits: Only shallow soil / rock geochemical data are needed to infer deep geothermal potential, without the need for large-scale drilling operations, which significantly reduces exploration costs and time. The generated prediction map is highly consistent with known fault zones and can directly guide the deployment of subsequent drilling projects.
[0021] (4) Highly practical for engineering: The asymmetric decision-making strategy can adjust the recall rate constraint according to the exploration needs, adapt to the risk preferences of different scenarios, and the prediction results are intuitive and easy to understand. It is easily recognized and applied by geological experts and has broad engineering promotion value. Attached Figure Description
[0022] Figure 1 The flowchart illustrates the steps of a deep geothermal resource prediction method based on mechanism-space-distribution fusion, as provided in this embodiment of the invention.
[0023] Figure 2 This is a flowchart of the overall architecture of the MSD-Fusion multidimensional heterogeneous fusion architecture of the present invention.
[0024] Figure 3 The figure shows a comparison of the effects of feature extraction in the spatial perception layer and dynamic screening in RFE; where (a) is the spatial distribution of the local singularity index based on KNN, and (b) is the curve of the change of the number of features and the model AUC value during the RFE iteration process.
[0025] Figure 4 A schematic diagram of anchoring the asymmetric decision boundary for recall.
[0026] Figure 5 This is a prediction map of the deep geothermal resource potential in the Lantian area; the red areas are high-potential target areas, and the geological verification is carried out by overlaying known fault lines, which visually presents the target area delineation effect.
[0027] Figure 6This chart compares the three models SVM, KNN, and MSD-Fusion on four key metrics: AUC, Recall, Accuracy, and F2.
[0028] Figure 7 The figure shows the importance of the top 20 features. It displays the top 20 input features that contribute the most to the prediction results of the SGAF-GATv2 model in the geothermal anomaly detection task and their importance scores (arranged from left to right in order of importance from high to low). Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0030] Please see Figure 1 , Figure 1 A flowchart illustrating the steps of a deep geothermal resource prediction method based on mechanism-space-distribution fusion, as provided in this embodiment of the invention, is as follows: S1. Obtain geochemical multi-element analysis data and basic geological data of the target area, perform limit correction and standardization preprocessing on the geochemical multi-element analysis data, construct label vectors based on known geothermal anomaly information in the basic geological data, and divide the training set and test set. S2. Construct a multi-dimensional heterogeneous fusion feature system that coordinates the mechanism-driven layer, spatial perception layer and data distribution layer. Using standardized geochemical multi-element analysis data as input, generate multi-dimensional heterogeneous fusion features that integrate geothermal mineralization mechanism, geochemical field spatial structure and data distribution law. S3. A recursive feature elimination algorithm is adopted, with random forest as the base learner and Gini importance as the evaluation index. Low-importance features are eliminated in each round, and the core geothermal geochemical fingerprint set is obtained through iterative screening. S4. A three-element consensus integrated prediction module is constructed based on gradient boosting tree, random forest and extreme random tree. The core geothermal geochemical fingerprint set is used as input and the label vector is used as supervision. An equal-weight soft voting strategy is adopted to fuse the prediction results of each model and output the joint prediction probability. S5. Based on the recall rate anchoring strategy, determine the asymmetric decision boundary. Based on the joint prediction probability and label vector on the training set, set the minimum recall rate constraint, optimize the optimal threshold, perform full-domain grid prediction on the target area, and output the geothermal resource potential prediction map and high-potential target area.
[0031] In one specific embodiment of this example, the data preparation and preprocessing process includes: ① Data Acquisition: Acquire geochemical multi-element analysis data and basic geological data of the target area. In this embodiment of the invention, the geochemical multi-element analysis data includes sampling point number, spatial coordinates (longitude and latitude), and concentration values of 48 trace elements; the basic geological data includes regional geological map, fault structure distribution map, and information on known geothermal anomalies (hot springs, geothermal wells).
[0032] ② Correction for data below the detection limit: For geochemical multi-element analysis data below the detection limit (LOD), a half-limit substitution method is used for correction. The correction formula is as follows: In the formula, This represents the detection limit for the j-th element. The string-formatted element value is first parsed and converted, then corrected according to the formula above.
[0033] ③ Standardization: To eliminate dimensional differences between different elements, Z-Score standardization is used. The formula is: In the formula, and The first j The mean and standard deviation of each element across all samples. After standardization, the mean concentration of each element is 0 and the variance is 1.
[0034] ④ Label Vector Construction: Label vectors are constructed based on known geothermal anomaly information in the basic geological data. Samples located in known geothermal fields, hot spring outcrops, or geothermal well locations are labeled as positive samples (y=1), and the remaining samples are labeled as negative samples (y=0). It is important to note that geothermal exploration is a typical scenario of "positive samples are known, negative samples are potential." Negative samples only indicate that no geothermal anomalies have been found so far, not that there is absolutely no resource potential.
[0035] ⑤ Dataset partitioning: A stratified random sampling method is used to partition the dataset into 80% training set and 20% test set to ensure that the ratio of positive to negative samples in the training and test sets is consistent with that in the original dataset. In this embodiment, the training set has 112 samples (approximately 33 positive samples and approximately 79 negative samples), and the test set has 29 samples (9 positive samples and 20 negative samples).
[0036] In one specific embodiment of this example, to address the core problems of existing geothermal exploration technologies, such as the lack of physical mechanisms, high underreporting rates, and weak generalization capabilities, such as... Figure 2As shown, this invention constructs a multi-dimensional heterogeneous fusion architecture based on "mechanism-space-distribution" (MSD-Fusion). This architecture consists of four key modules: a multi-dimensional heterogeneous fusion feature module, a dynamic progressive feature denoising module, a ternary consensus integrated prediction module, and a recall-anchored asymmetric decision-making module. Specifically, the multi-dimensional heterogeneous fusion feature module constructs a feature system that incorporates geological mechanisms, spatial topology, and data distribution patterns, providing high-value input for the prediction model; the dynamic progressive feature denoising module removes redundant noise features and filters core geochemical fingerprints; the ternary consensus integrated prediction module integrates the advantages of multiple models to improve prediction stability; and the recall-anchored asymmetric decision-making module adapts to the risk preferences of geothermal exploration to achieve the goal of zero missed reports. Figure 2 As shown, the input of MSD-Fusion is geochemical multi-element analysis data and basic geological data of the target area. After preprocessing, it forms a standardized data matrix and label vector. The output is a geothermal resource potential prediction map and high-potential target areas, which can be directly used to guide the deployment of drilling projects.
[0037] In one specific implementation of this embodiment, the process of constructing a multidimensional heterogeneous fusion feature system includes: In geothermal prediction research, existing geothermal prediction methods often employ flattened feature inputs or focus solely on one dimension such as mechanism, space, or data, resulting in insufficient feature representation capabilities and difficulty in adapting to complex geological scenarios. For example, traditional geostatistical methods rely solely on single-element concentration features, neglecting the spatial correlation of hydrothermal migration and mineralization mechanisms; conventional machine learning methods lack geological prior constraints, are prone to learning spurious correlations, and have poor model interpretability.
[0038] To address the aforementioned limitations, this invention proposes a multi-dimensional heterogeneous fusion feature module. This module constructs a three-layer collaborative architecture comprising a mechanism-driven layer, a spatial perception layer, and a data distribution layer, explicitly fusing geothermal mineralization mechanisms, the spatial structure of geochemical fields, and data distribution patterns. For example... Figure 2 As shown, this module stitches and fuses the three layers of features to form a multi-dimensional feature matrix that combines geological plausibility with data discriminative power, providing a solid foundation for subsequent predictions. The specific construction of each layer is as follows: (1) Mechanism-driven layer construction Based on the theory of radiothermal heat generation and the mechanism of hydrothermal convection, the mechanism-driven layer takes standardized elemental concentration data as input and outputs three core mechanism factors, forming a 3-dimensional mechanism feature vector. The operation steps are as follows: ① Calculation of radioactive heat generation potential: Based on the decay heat generation characteristics of the three core radioactive elements, uranium (U), thorium (Th), and potassium (K), the heat generation potential index is calculated: In the formula, , , These represent the elemental weights of uranium, thorium, and potassium, respectively. , , These represent the standardized concentrations of uranium, thorium, and potassium, respectively. The heat generation potential index characterizes the core heat source potential for geothermal generation in the region.
[0039] ② Construction of Deep Fluid Tracer Index: Based on the enrichment patterns of elements in the leading edge halo of hydrothermal activity, seven fluid indicator elements—arsenic (As), antimony (Sb), lithium (Li), rubidium (Rb), cesium (Cs), tungsten (W), and tin (Sn)—were aggregated. The formula is as follows: In the formula, the concentrations of each element are standardized values. This index characterizes the channels and ranges of hydrothermal fluid migration, providing key clues for locating geothermal anomalies.
[0040] ③ Generation of heat-mass migration coupling factor: Construct a nonlinear interaction term to characterize the "heat source-channel" synergistic mechanism, the formula is: This factor accurately captures the core geological mechanisms of geothermal formation, achieving deep coupling between heat source potential and fluid channel information.
[0041] (2) Construction of the spatial perception layer Geochemical fields exhibit significant spatial continuity and heterogeneity. Traditional methods neglect spatial topological relationships, treating sampling points as independent samples, leading to difficulties in capturing weak anomaly signals. Therefore, the spatial sensing layer constructed in this embodiment of the invention uses the spatial coordinates of sampling points and the original elemental concentration data (unstandardized) as input. It constructs a spatial topological structure based on a K-nearest neighbor graph, employs the Ball-Tree algorithm to efficiently search for nearest neighbors, and outputs multi-scale spatial features. The operation steps are as follows: ① Multi-scale spatial lag feature extraction: Two differentiated scales, K=5 (local micro field) and K=15 (regional background field), are set. The mean and variance of element concentration in the corresponding neighborhood of each sampling point are calculated. The regional background trend and local anomalous fluctuations of the geochemical field are captured simultaneously to form 4-dimensional spatial features.
[0042] ② Calculation of Fractal Singularity Index: Based on the concept of fractal filtering, weak anomalous signals are identified. The formula is as follows: In the formula, Indicates the element concentration at the sampling point. This represents the neighborhood mean at the corresponding scale. Represents the smoothing term and (To avoid the denominator being zero). Indicates a positive anomaly (element enrichment). Indicates a negative anomaly (element depletion). For example... Figure 3 As shown in (a), the fractal singularity index can identify weak but regular mineral-induced anomalous signals from strong background interference, achieving accurate capture of weak anomalous signals and forming 1-dimensional singularity features. Finally, the spatial perception layer outputs a 5-dimensional spatial feature vector.
[0043] (3) Construction of data distribution layer Geothermal exploration falls into the category of "positive samples are known, negative samples are potential," where the ambiguous definition of negative samples easily leads to label noise. Therefore, the data distribution layer uses standardized elemental concentration data as input and employs the Isolation Forest algorithm for unsupervised distribution learning, outputting distribution anomaly scores to reduce reliance on label data. The operation steps are as follows: Isolation Forest algorithm parameter settings: 100 trees, 25 sample splits, and contamination level set to "auto" (automatically estimates the proportion of outliers in the dataset). After training, the output decision function value is inverted, so that a higher score represents a stronger degree of anomaly, forming a 1D data distribution feature vector. The anomaly score serves as a prior probability, deeply complementing the supervised learning features. Its value lies in the fact that even if a region currently lacks known geothermal hotspots, as long as its element combination features significantly deviate from the normal field in statistical distribution, the model will still pay attention to this feature, effectively suppressing label noise interference.
[0044] (4) Feature fusion By concatenating and fusing the 3D features of the mechanism-driven layer, the 5D features of the spatial perception layer, and the 1D features of the data distribution layer, a 9D heterogeneous fusion feature matrix is obtained: This feature matrix combines geological rationality (mechanistic layer), spatial sensitivity (spatial layer), and data intuition (distribution layer), providing high-quality input for subsequent predictions.
[0045] In one specific implementation of this embodiment, the dynamic progressive feature denoising process includes: Geochemical data typically contains hundreds of elemental features, exhibiting high-dimensional redundancy and noise interference. Directly inputting these data into models can easily lead to overfitting and reduced generalization ability. Traditional feature selection methods, such as Principal Component Analysis (PCA), are constrained by linear assumptions and struggle to adapt to the feature distribution of nonlinear geological data. Therefore, this invention employs a Recursive Feature Elimination (RFE) algorithm to construct a dynamic, progressive feature denoising module, using a random forest as the base learner to iteratively select core features through multiple rounds. Specific steps include: (1) Initialization The base model is set as a random forest classifier, and the initial feature set is a combination of 9-dimensional heterogeneous fusion features and 342-dimensional original element features, for a total of 351 features.
[0046] (2) Iterative filtering and output After each training round, the Gini importance of each feature is calculated, and the bottom 5% of features are removed. The feature set is then updated, and training is repeated. This iterative process continues until the number of features drops to a preset target of 50 dimensions. After iteration, the optimal 50-dimensional core geothermal geochemical fingerprint set is obtained. High-value features automatically retained during the screening process include: thermal-mass migration coupling factor, tungsten spatial singularity index, and isolated forest anomaly score. For example... Figure 3 (b) shows the curves of the change in the number of features and the model AUC value during the RFE iteration process, which verifies that the model performance is optimal when the number of features is reduced to 50.
[0047] In one specific implementation of this embodiment, the ternary consensus integration prediction process includes: Single models are prone to excessive variance and insufficient generalization ability in small-sample, high-noise geological data, making it difficult to balance prediction accuracy and stability. For example, Gradient Boosting Tree (GBDT) is prone to overfitting, Random Forest has limited noise resistance, and Extremely Random Tree is insufficient in capturing local features. Therefore, this invention integrates three heterogeneous tree models—GBDT, Random Forest, and Extremely Random Tree—to construct a ternary consensus integrated prediction module. A soft voting strategy is used to combine the advantages of each model. The operation steps are as follows: (1) Base model training Based on core geothermal geochemical fingerprints Three heterogeneous tree models were trained separately, with the following parameter configurations for each model: Table 1 Parameter Configuration Table for Different Heterogeneous Tree Models All models have a random seed of 42 to ensure reproducible results. Among them, GBDT is good at capturing nonlinear patterns to reduce bias, random forest reduces variance through randomness, and extreme random trees further enhance noise resistance. The three models complement each other.
[0048] (2) Soft voting consensus decision Using an equal-weighted soft voting strategy, the joint prediction probability of the three models for the sample belonging to the "geothermal anomaly" category is calculated: In the formula, x This represents the input sample feature vector. y =1 indicates that the sample belongs to the anomaly class. , , These represent the probability values that GBDT, Random Forest, and Extremely Random Tree predict a feature x as belonging to the positive class (y = 1), respectively. The final score increases significantly only when all three models give high-confidence predictions, effectively suppressing random misjudgments by a single model and improving prediction stability and reliability.
[0049] In one specific implementation of this embodiment, the recall rate-anchored asymmetric decision-making process is as follows: The core risk preference in geothermal exploration is "better to misjudge than to miss," as the cost of underreporting far outweighs the cost of false positives. Traditional symmetric decision boundaries are ill-suited to this requirement, easily leading to high underreporting rates. Therefore, this invention proposes a recall-anchored asymmetric decision-making strategy, optimizing the decision boundary by fixing a minimum recall constraint. The operational steps are as follows: (1) Optimal threshold optimization Joint prediction probability on the training set Using the true label vector as input, set a minimum recall constraint. A grid search is used to traverse all candidate thresholds within the range of 0.1 to 0.9 with a step size of 0.01. Recall and precision are calculated for each threshold, and the threshold corresponding to the maximum precision that satisfies the recall constraint is found. ,like Figure 4 This demonstrates the threshold cutoff point and decision boundary distribution corresponding to the optimal precision under a fixed recall constraint. If geothermal exploration requires zero false negatives, the minimum recall constraint can be adjusted to... If the preset recall constraint cannot be met after iterating through all thresholds, the system will automatically switch to the F2 score maximization strategy to determine the threshold. In the F2 score, the adjustment factor β=2 makes the recall weight twice that of the precision, adapting to different risk preferences in different scenarios.
[0050] (2) Global prediction and mapping Using the standardized feature data of all grid nodes in the study area as input, the probability field is obtained through steps S2 to S4. With the optimal threshold Perform binarization determination ( hour, ;otherwise It outputs a 50m×50m resolution geothermal resource potential prediction map and high-potential target areas. The contrast of the visualization color markers is optimized through histogram equalization to clearly distinguish the background area, high-potential prediction area, known training wells, and validation geothermal wells.
[0051] (3) Geological verification The predicted results were overlaid and compared with known fault zones in the target area to verify that the location of high-potential target areas is consistent with geological patterns. In this embodiment, the predicted high-potential target areas highly coincide with the main fault zones in the region, verifying the geological rationality of the model's prediction results.
[0052] Experimental Results and Analysis: This embodiment verifies the effectiveness of the method of the present invention on a geothermal exploration dataset from the Lantian area. This dataset contains the spatial coordinates of 141 sampling points and the concentration data of 48 trace elements (approximately 56KB), including 42 geothermal anomalies and 99 non-anomalies. Stratified random sampling was used to divide the dataset into an 80% training set (112 samples) and a 20% test set (29 samples), ensuring strict isolation between training and test data. To comprehensively evaluate model performance, the present invention uses the following evaluation metrics: Area Under the ROC Curve (AUC), Recall, Accuracy, and F2-Score. Recall is used as the core metric to meet the requirement of zero false negatives in geothermal exploration; F2-Score considers both recall and precision.
[0053] Experimental results on independent test sets demonstrate that the proposed multi-source heterogeneous data fusion method (MSD-Fusion) achieves significantly better performance than traditional methods. Specific experimental data are shown in Table 2.
[0054] Table 2 shows the performance comparison of different models on the Lantian dataset. Table 2 shows the experimental results: (1) Zero false negatives: The method of this invention achieved a recall rate of 1.0000 on the test set, successfully identifying all potential geothermal anomalies without any omissions. In contrast, the recall rates of traditional SVM and KNN methods are 71.42% and 85.71%, respectively, which have a greater risk of false negatives.
[0055] (2) Improved overall performance: The AUC of the method of this invention is as high as 0.9805, which is about 3.89% higher than the SVM model and about 4.22% higher than the KNN model, indicating that the model has a significant advantage in the overall ability to distinguish between positive and negative samples.
[0056] (3) Engineering application value: Although the accuracy rate is slightly reduced (0.7931) in pursuit of 100% recall rate, it meets the actual needs of "better to be broad than to miss" in the initial geothermal exploration stage - delineate the target area as comprehensively as possible through low-cost calculation methods, and then eliminate false alarms through a small amount of subsequent drilling.
[0057] Figure 5 A map showing the predicted potential of deep geothermal resources in the Lantian area was presented. The red high-potential target areas are highly consistent with known fault zones, verifying the geological rationality of the prediction results. Figure 6The performance comparison of three models, SVM, KNN, and MSD-Fusion, is presented. Figure 7 The ranking of the top 20 features by importance is shown. High-value features such as thermal-mass migration coupling factor, tungsten element spatial singularity index, and isolated forest anomaly score are automatically retained, demonstrating the effectiveness of the multidimensional heterogeneous fusion feature system.
[0058] In one specific implementation of this embodiment, the deployment and application of the model are as follows: (1) Model workpiece storage and application scenarios This invention utilizes Joblib to store the complete artifact, consisting of the trained ensemble model, normalizer, feature selector, and optimal threshold, supporting direct loading and deployment without requiring repeated training. Furthermore, this invention is applicable to scenarios such as deep geothermal resource exploration, dry hot rock target area selection, hydrothermal geothermal field boundary delineation, and pre-exploitation assessment of geothermal resources. It enables the prediction of deep, concealed geothermal resources using shallow surface soil or rock geochemical data, significantly reducing exploration costs and time.
[0059] (2) Visual output This invention also includes a full-process visualization module, which can automatically generate four core diagrams: an architecture flowchart showing the complete technical route from data preprocessing to target output (…). Figure 2 The spatial singularity and RFE curve comparison of feature extraction and filtering effects are presented intuitively. Figure 3 This illustrates a schematic diagram of recall-anchored decision boundaries in the asymmetric decision threshold optimization process. Figure 4 Geothermal resource potential prediction map (which delineates high-potential target areas to guide engineering deployment) Figure 5 ).
[0060] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions conceived without inventive effort should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims.
[0061] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for predicting deep geothermal resources based on mechanism-spatial-distribution fusion, characterized in that, include: S1. Obtain geochemical multi-element analysis data and basic geological data of the target area, perform limit correction and standardization preprocessing on the geochemical multi-element analysis data, construct label vectors based on known geothermal anomaly information in the basic geological data, and divide the training set and test set. S2. Construct a multi-dimensional heterogeneous fusion feature system that coordinates the mechanism-driven layer, spatial perception layer and data distribution layer. Using standardized geochemical multi-element analysis data as input, generate multi-dimensional heterogeneous fusion features that integrate geothermal mineralization mechanism, geochemical field spatial structure and data distribution law. S3. A recursive feature elimination algorithm is adopted, with random forest as the base learner and Gini importance as the evaluation index. Low-importance features are eliminated in each round, and the core geothermal geochemical fingerprint set is obtained through iterative screening. S4. A three-element consensus integrated prediction module is constructed based on gradient boosting tree, random forest and extreme random tree. The core geothermal geochemical fingerprint set is used as input and the label vector is used as supervision. An equal-weight soft voting strategy is adopted to fuse the prediction results of each model and output the joint prediction probability. S5. Based on the recall rate anchoring strategy, determine the asymmetric decision boundary. Based on the joint prediction probability and label vector on the training set, set the minimum recall rate constraint, optimize the optimal threshold, perform full-domain grid prediction on the target area, and output the geothermal resource potential prediction map and high-potential target area.
2. The method according to claim 1, characterized in that, In step S2, the mechanism-driven layer takes standardized elemental concentration data as input and outputs radioactive heat generation potential energy. Hydrothermal fluid tracer index and the thermal-mass migration coupling factor generated by the nonlinear interaction between the two. This constitutes a 3D mechanism feature vector; wherein, the radioactive heat generation potential energy Based on the decay heat generation characteristics of uranium, thorium, and potassium, the formula is: In the formula, , , These represent the elemental weights of uranium, thorium, and potassium, respectively. , , These represent the concentrations of uranium, thorium, and potassium, respectively; the hydrothermal fluid tracer index... The formula for polymerizing seven fluid indicator elements—arsenic, antimony, lithium, rubidium, cesium, tungsten, and tin—is as follows: The thermal-mass migration coupling factor The product of the heat generation potential energy and the fluid tracer index is given by the formula: .
3. The method according to claim 1, characterized in that, In step S2, the spatial perception layer takes the spatial coordinates of the sampling points and elemental concentration data as input, constructs a spatial topology based on the K-nearest neighbor graph, and outputs the neighborhood mean and variance at two scales: K=5 and K=15, as well as the fractal singularity index used to identify weak anomalous signals. This forms a 5-dimensional spatial feature vector; the formula for calculating the fractal singularity index is: In the formula, Indicates the element concentration at the sampling point. This represents the neighborhood mean at the corresponding scale. This indicates the smoothing term.
4. The method according to claim 1, characterized in that, In step S2, the data distribution layer takes standardized element concentration data as input and uses the isolated forest algorithm for unsupervised training, outputting distribution anomaly scores to form a 1-dimensional data distribution feature vector. The isolated forest algorithm is set to 100 trees, 25 sample splits, and "auto" pollution level. The algorithm output decision function value is inverted so that a higher score represents a stronger degree of anomaly, which complements the prior probability and supervised learning features.
5. The method according to claim 1, characterized in that, In step S3, the input of the recursive feature elimination algorithm is a multidimensional initial feature set composed of multidimensional heterogeneous fusion features output from the mechanism-driven layer, spatial perception layer, and data distribution layer, and multidimensional original element features. After each round of training, the Gini importance is calculated and the features ranked in the bottom 5% are removed. The feature set is iteratively updated until a 50-dimensional core geothermal geochemical fingerprint set is retained.
6. The method according to claim 1, characterized in that, In step S4, the parameters of the gradient boosting tree model are configured as follows: learning rate = 0.05, number of trees = 100, maximum depth = 3, minimum number of leaf node samples = 3, subsample ratio = 0.8; the parameters of the random forest model are configured as follows: number of trees = 150, maximum depth = 5, minimum number of leaf node samples = 2, and class weight balance; the parameters of the extreme random tree model are configured as follows: number of trees = 150, maximum depth = 10, minimum number of leaf node samples = 2, random feature splitting, and class weight balance; the random seed for all models is set to 42.
7. The method according to claim 1, characterized in that, In step S5, the recall anchoring strategy takes the joint predicted probability and label vector on the training set as input and sets a minimum recall constraint. Or adjust to meet the requirement of zero underreporting in geothermal exploration. The algorithm iterates through candidate thresholds in the range of 0.1 to 0.9 using a grid search with a step size of 0.01, outputting the optimal threshold that satisfies the recall constraint. If the preset recall constraint cannot be met after iteration, the algorithm automatically switches to the F2 score maximization strategy to determine the threshold, where the F2 score includes an adjustment factor. β =2.
8. The method according to claim 1, characterized in that, In step S5, the global prediction takes the standardized feature data of all grid nodes in the study area as input, and obtains the probability field after processing in steps S2-S4. The optimal threshold is used for binarization judgment, and a geothermal resource potential prediction map and high potential target area with a resolution of 50m×50m are output. The contrast of the visualization color mark is optimized by histogram equalization, and the prediction results are superimposed and compared with the known fault tectonic zones in the target area for verification.
9. The method according to claim 1, characterized in that, The method is applicable to scenarios such as deep geothermal resource exploration, dry hot rock target area selection, hydrothermal geothermal field boundary delineation, or pre-exploitation assessment of geothermal resources. It enables the prediction of deep hidden geothermal resources based on shallow surface soil or rock geochemical data. The method also includes a model deployment process, which uses Joblib to save the complete artifact consisting of the trained ensemble model, normalizer, feature selector, and optimal threshold, and supports direct loading and deployment.
10. The method according to claim 1, characterized in that, The method also includes a full-process visualization module, which automatically generates an architecture flowchart, a spatial singularity and RFE curve comparison chart, a recall rate anchoring decision boundary diagram, and a geothermal resource potential prediction map for technical process display and prediction result output.