Small sample geothermal resource prediction method, device and system, and storage medium
By using a small-sample geothermal resource prediction method, hierarchical category sampling and Bayesian optimization algorithm are used to generate homogeneous synthesis tasks, and a multi-scale Bayesian prototype is constructed. This solves the problem of limited sample size in geothermal resource exploration and achieves high-precision, stable and reliable prediction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies in geothermal resource exploration suffer from problems such as limited sample size, high data acquisition costs, reliance on expert experience, and poor repeatability, resulting in insufficient accuracy and stability in geothermal resource evaluation.
A small-sample geothermal resource prediction method is adopted. By generating homogeneous synthesis tasks through hierarchical category sampling and adaptive Gaussian noise, the highest priority hyperparameter is automatically searched by a Bayesian optimization algorithm based on Gaussian process. A Bayesian prototype is constructed at multiple scales, and similarity is calculated and mixed with Student distribution for prediction.
It significantly improves the accuracy and stability of geothermal resource prediction under small sample conditions, reduces prediction error and variance, provides reliable confidence intervals to support decision-making, reduces the cost of manual intervention, and improves computational efficiency.
Smart Images

Figure CN121524966B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of geothermal resource exploration technology, specifically relating to a method, device, system, and storage medium for predicting small-sample geothermal resources. Background Technology
[0002] Geothermal resource assessment is a core component of energy exploration. Currently, the mainstream resource assessment methods mainly include:
[0003] 1. Geological analogy method: It relies heavily on the personal experience of experts, is highly subjective, and the conclusions of different experts vary greatly, resulting in poor repeatability and generalizability.
[0004] 2. Numerical simulation method: It requires extremely detailed geological structure and fluid dynamic parameters, and the data acquisition cost is high and the cycle is long, making it difficult to apply in the early stages of actual regional exploration.
[0005] 3. Traditional machine learning methods, such as neural networks and random forests, perform well when there is sufficient data, but their performance is heavily dependent on a large number of labeled samples (usually more than 500), while the actual available samples in geothermal exploration are usually extremely limited. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a method, device, system, and storage medium for predicting geothermal resources with a small sample size. Under the objective constraints of small sample size, high cost, and strong regional heterogeneity in geothermal resource exploration, this invention achieves high-precision, high-robustness, and quantitative uncertainty-free reserve prediction.
[0007] To achieve the above objectives, the present invention provides the following solution:
[0008] A method for predicting geothermal resources in a small sample size includes:
[0009] Step S1: From the source dataset containing historical data of multiple geothermal regions, generate multiple homogeneous synthesis tasks by hierarchical category sampling and adding adaptive Gaussian noise;
[0010] Step S2: In the homology synthesis task, a Bayesian optimization algorithm based on Gaussian process is used to automatically search for the highest priority hyperparameter of the Bayesian prior distribution.
[0011] Step S3: On a small sample of data in a new target region, fix the highest priority hyperparameter and only fine-tune the sharpening parameter, temperature parameter and upper limit of sample weight;
[0012] Step S4: Discretize the continuous geological features in the small sample data of the target area at three scales: coarse, medium, and fine. Based on the normal-inverse Wishart conjugate prior, construct a Bayesian prototype for each geothermal resource category at each scale.
[0013] Step S5: For the sample to be predicted, calculate its similarity with historical samples at coarse, medium and fine scales using the Bayesian prototype. The coarse scale uses the discrete rank difference measure, the medium scale uses a mixture of discrete rank and continuous feature posterior density measure, and the fine scale uses the posterior prediction density measure.
[0014] Step S6: Merge the similarity of each scale according to the preset weight, and after sharpening and upper limit truncation, normalize to obtain the weight of each sample.
[0015] Step S7: Aggregate the category weights based on the sample weights, and mix the Bayesian prediction distributions of each category into a mixed student. The distribution is calculated, and the predicted mean and confidence intervals at the specified confidence level are output.
[0016] The present invention also provides a small-sample geothermal resource prediction device, comprising:
[0017] The first processing module is used to generate multiple homogeneous synthesis tasks from a source dataset containing historical data of multiple geothermal regions by hierarchical category sampling and adding adaptive Gaussian noise.
[0018] The second processing module is used to automatically search for the highest priority hyperparameter of the Bayesian prior distribution in the homogeneous synthesis task by employing a Bayesian optimization algorithm based on Gaussian processes.
[0019] The third processing module is used to fix the highest priority hyperparameter on a small sample of data in a new target area, and only fine-tune the sharpening parameter, temperature parameter and upper limit of sample weight.
[0020] The fourth processing module is used to discretize the continuous geological features in the small sample data of the target area at three scales: coarse, medium and fine, and to construct a Bayesian prototype for each geothermal resource category at each scale based on the normal-inverse Wishart conjugate prior.
[0021] The fifth processing module is used to calculate the similarity between the sample to be predicted and historical samples at coarse, medium and fine scales using the Bayesian prototype. The coarse scale uses the discrete rank difference measure, the medium scale uses a hybrid measure of discrete rank and continuous feature posterior density, and the fine scale uses the posterior prediction density measure.
[0022] The sixth processing module is used to fuse the similarity of each scale according to a preset weight, and after sharpening and upper limit truncation, normalize to obtain the weight of each sample.
[0023] The seventh processing module is used to aggregate class weights based on sample weights and mix the Bayesian prediction distributions of each class into a mixed student model. The distribution is calculated, and the predicted mean and confidence intervals at the specified confidence level are output.
[0024] The present invention also provides a small sample geothermal resource prediction system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a small sample geothermal resource prediction method when run by the processor.
[0025] The present invention also provides a storage medium storing a computer program, which executes a small-sample geothermal resource prediction method when running.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] 1. Meta-learning-enhanced automatic optimization of Bayesian priors: The meta-learning framework is applied to the field of geothermal resource evaluation. By learning from synthetic tasks, the optimal Bayesian prior hyperparameters are automatically obtained, which solves the bottleneck problem of difficult prior setting under small sample conditions.
[0028] 2. Multi-scale differentiated similarity calculation strategy: A strategy of selecting different similarity measures according to scale characteristics is proposed, which takes into account the robustness of coarse scale, the balance of meso scale and the accuracy of fine scale, and significantly improves the accuracy of analogy prediction.
[0029] 3. Based on mixed student groups Quantification of distribution uncertainty: A rigorous Bayesian mixture model output format was constructed, which not only provides point predictions but also reliable confidence intervals, greatly enhancing the decision support value of the prediction results. Attached Figure Description
[0030] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart of the small-sample geothermal resource prediction method according to an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] Example 1
[0035] like Figure 1 As shown, this invention provides a method for predicting geothermal resources in small samples, comprising:
[0036] Step S1: From the source dataset containing historical data of multiple geothermal regions, generate multiple homogeneous synthesis tasks by hierarchical category sampling and adding adaptive Gaussian noise;
[0037] Step S2: In the homology synthesis task, a Bayesian optimization algorithm based on Gaussian process is used to automatically search for the highest priority hyperparameter of the Bayesian prior distribution.
[0038] Step S3: On a small sample of data in a new target region, fix the highest priority hyperparameter and only fine-tune the sharpening parameter, temperature parameter and upper limit of sample weight;
[0039] Step S4: Discretize the continuous geological features in the small sample data of the target area at three scales: coarse, medium, and fine. Based on the normal-inverse Wishart conjugate prior, construct a Bayesian prototype for each geothermal resource category at each scale.
[0040] Step S5: For the sample to be predicted, calculate its similarity with historical samples at coarse, medium, and fine scales using a Bayesian prototype. The coarse scale uses a discrete rank difference metric, the medium scale uses a mixture of discrete rank and continuous feature posterior density, and the fine scale uses a posterior prediction density metric. The coarse-scale discrete rank difference metric uses an exponential kernel function based on Manhattan distance. The medium-scale mixture is a weighted average of discrete rank similarity and continuous feature posterior density similarity. The fine-scale posterior prediction density metric is based on a multivariate student... The log probability density of the distribution;
[0041] Step S6: Merge the similarity of each scale according to the preset weight, and after sharpening and upper limit truncation, normalize to obtain the weight of each sample.
[0042] Step S7: Aggregate the category weights based on the sample weights, and mix the Bayesian prediction distributions of each category into a mixed student. The distribution is calculated, and the predicted mean and confidence intervals at the specified confidence level are output.
[0043] As one embodiment of the present invention, step S1 includes:
[0044] Step S11: For each geothermal resource category in the source dataset, perform stratified sampling. One sample is used as the prototype center;
[0045] Step S12: For each consecutive feature of each prototype center Calculate its adaptive noise level ;
[0046] Step S13: Add Gaussian noise to each prototype center. generate A new sample is created, inheriting the category label from the prototype;
[0047] Step S14: Repeat steps S11-S13. Next, generate containing A set of synthetic tasks.
[0048] Furthermore, in step S11, the source dataset is: ,in, For feature vectors, For the target value, Category labels;
[0049] For each category Calculate the number of prototypes for each category. :
[0050] ;
[0051] in, The number of synthesis tasks. Number of samples for each synthesis task; For category The number of samples in the source dataset;
[0052] From category Random sampling of data The prototype is: ;
[0053] Furthermore, in step S12, the standard deviation of the features of the source dataset is calculated as follows:
[0054] ;
[0055] in, For the source dataset,
[0056] Calculate the adaptive noise level: ,in, This represents the noise figure.
[0057] Furthermore, in step S13, for each prototype ,in, The union of all categories, For category The prototype set;
[0058] Calculate the number of samples needed to generate for each prototype: ;
[0059] generate One perturbation sample:
[0060] ;
[0061] ;
[0062] in, To generate sample feature vectors, the dimension ; It is a Gaussian noise vector; As prototype The corresponding target value; The standard deviation of all target values in the source dataset; The target value for the generated k-th perturbation sample; This is random noise added to the target value.
[0063] Furthermore, in step S14, steps 11 to 13 are repeated. Next, generate Each task is independent; the output is a composite task set. Each task contains One sample: ;in, Indicates a containing A collection of n elements, each element They are all part of a synthesis task; To generate the feature vector of the sample; The target value for generating the sample.
[0064] As one embodiment of the present invention, in step S2, the feature prior accuracy is defined. Degrees of freedom Covariance Scale Target prior parameters Temperature parameters Sharpening parameters The search space is defined; a Gaussian process regression model is initialized as a surrogate model; candidate hyperparameters to be evaluated are selected using an expected improved acquisition function; the performance of the candidate points is evaluated by cross-validation on a synthetic task set; the surrogate model is iteratively updated and new candidate points are selected until the convergence condition is met or the maximum number of iterations is reached, to obtain the highest priority a priori hyperparameters of the Bayesian prior distribution. Specifically, this includes:
[0065] Step 21: Define the hyperparameter search space
[0066] Define an eight-dimensional hyperparameter space : ;
[0067] Step 22: Initialize the Gaussian process model
[0068] Using the Matern kernel function:
[0069] ;
[0070] in, Let be two points in the hyperparameter space; For the second type of modified Bessel function; set the smoothing parameter. , length scale ;
[0071] Step 23: Randomly initialize sampling
[0072] Random sampling Initial point
[0073] For each Calculate its average loss across all synthesis tasks: ,in, Indicates the use of hyperparameters The algorithm, It is an abbreviation for Algorithm.
[0074] Step 24: Bayesian optimization of the main loop
[0075] For each iteration arrive :
[0076] 1. Update the Gaussian process posterior: based on the current observation dataset. The kernel matrix and its inverse of the Gaussian process model are recalculated to obtain information about the loss function. Complete posterior distribution This distribution describes arbitrary hyperparameters based on the tested points. The probability estimate of the corresponding loss.
[0077] 2. Calculate the expected improvement: Use the expected improvement acquisition function: ,in This is the best loss value observed so far. The formula calculates the new point. How much improvement is expected in the loss value compared to the current best value?
[0078] 3. Select the next candidate point: Optimize the hyperparameter space using an algorithm (such as L-BFGS). Searching for what can make The largest point, that is
[0079] .
[0080] 4. Evaluate candidate points: Combine the selected hyperparameters. All generated in step S1 Cross-validation was performed on each synthetic task to calculate its average loss.
[0081] 5. Update the observation dataset: Add the newly evaluated points and their loss values to the historical records to form a new dataset. .
[0082] Step 25: Early Stop Judgment
[0083] If continuous If no improvement is made in the next iteration, the process will terminate early.
[0084] Step 26: Return the highest priority hyperparameter.
[0085] ;
[0086] In one embodiment of the present invention, in step S3, the adaptation objective is to minimize the prediction loss on the target region dataset, and the gradient update is as follows:
[0087] ;
[0088] in, , To adapt to the temperature parameters, To adapt to the post-sharpening parameters, To accommodate the upper limit of weight after adaptation, For a small sample dataset of the target region, η This is the learning rate.
[0089] In one embodiment of the present invention, in step S4, the number of discretization levels for the coarse, medium, and fine scales are 3-5, 6-10, and 11-15, respectively.
[0090] Furthermore, in step S4, based on the normal-inverse Wishart conjugate prior, a Bayesian prototype is constructed for each geothermal resource category at each scale. Specifically, this includes: assuming a continuous geological feature vector for each category. Follows a multivariate normal distribution Using the normal-inverse Wishart distribution as parameters Conjugate priors:
[0091] ;
[0092] in, and These are standard notations, representing the multivariate normal distribution and the inverse Wishart distribution, respectively; For the mean Prior information; To measure the prior mean The confidence strength can be regarded as the "prior virtual sample size"; For the covariance matrix The prior information is a scale matrix; represents the degrees of freedom of the inverse Wishart distribution.
[0093] For a category Sample Its posterior distribution It still follows a normal-inverse Wishart distribution, with the parameters updated as follows:
[0094] ;
[0095] ;
[0096] ;
[0097] ;
[0098] in, It is the posterior mean precision weight. It is the posterior mean vector These are the posterior degrees of freedom. It is the sample mean. It is the sample scatter matrix;
[0099] Sample mean: ;
[0100] Sample scatter matrix: .
[0101] This posterior distribution is the Bayesian prototype of the category, which fully describes the understanding of the category feature distribution after observing the data.
[0102] As one embodiment of the present invention, in step S5...
[0103] Coarse-scale similarity is a similarity measure based on discrete level differences, i.e.:
[0104] ;
[0105] Mesoscale similarity is a hybrid of discrete-level similarity and continuous-feature posterior density similarity, i.e.:
[0106] ;
[0107] Fine-scale similarity is based on multiple students The posterior predictive density of the distribution is:
[0108] ;
[0109] in, For coarse-scale similarity, For mesoscale similarity, For fine-scale similarity, For coarse-scale temperature parameters, For feature weights, The discrete levels of the sample to be predicted, For the discrete levels of historical samples, For mixed weights, Let be the feature vector of the sample to be predicted. Let be the posterior mean vector. For posterior parameters, is the posterior scaling matrix.
[0110] In one embodiment of the present invention, in step S6...
[0111] Weight calculation formula:
[0112] Original similarity:
[0113] ;
[0114] Sharpening and Truncation:
[0115] ;
[0116] ;
[0117] Normalized weights:
[0118] ;
[0119] in, For the sample The original similarity, For the sample The similarity after sharpening For the sample Truncated similarity For the sample The final normalized weights, For scale The fusion weight, This is the upper limit parameter for weights. This refers to the sharpening index.
[0120] In one embodiment of the present invention, in step S7, the weights of each category are aggregated based on the normalized sample weights; the Bayesian prediction distributions of each category (student) are then processed. The distributions are weighted and mixed to form the final mixed student. Distribution; output the predicted mean (point prediction) and the specified confidence interval (uncertainty quantification) from the distribution.
[0121] Mixed students distributed:
[0122] ;
[0123] Predicted mean (point prediction):
[0124] ;
[0125] Prediction variance:
[0126] ;
[0127] Effective degrees of freedom (approximate):
[0128] ;
[0129] Confidence interval:
[0130] ;
[0131] in, For category Aggregate weights, ; For category The posterior predicted mean (for the target variable); For category The posterior prediction variance (for the target variable); For category The posterior degrees of freedom (for the target variable); This is the final predicted value (point prediction); For prediction variance (conditional variance); The confidence interval; For degrees of freedom of Distribution Quantiles.
[0132] This invention has the following characteristics:
[0133] 1. Significantly improves prediction accuracy and stability for small samples. In small sample scenarios with 20-50 samples, the prediction mean error (RMSE) is reduced by 20%-40% compared to traditional machine learning methods and by more than 20% compared to the standard Bayesian prototype method. The variance of the prediction results is reduced by more than 30%, and the stability is significantly enhanced.
[0134] 2. Achieve automated hyperparameter handling and efficient knowledge transfer. Eliminate the need for tedious manual parameter tuning by geological experts; automatically acquire prior hyperparameters with strong generalization capabilities through meta-learning. When exploring new areas, rapid adaptation can be completed in just 10-20 iterations, reducing manual intervention costs by over 80%.
[0135] 3. Multi-scale similarity metrics exhibit strong robustness and wide applicability. Coarse-scale similarity improves robustness by over 35% even with high feature noise; fine-scale similarity improves prediction accuracy by 25% with high-quality data. The model can adaptively learn the optimal weight combination for each scale.
[0136] 4. Provides reliable decision support. It outputs statistically significant confidence intervals (e.g., 90% interval coverage is increased from 75% in traditional methods to over 92%), enabling decision-makers to assess prediction risks. It also provides Top-K most similar historical samples, enhancing the model's interpretability.
[0137] 5. High computational efficiency and practical engineering applicability. The computationally intensive meta-learning phase can be performed offline, and the training can be reused for multiple target regions after a single training session. The fast adaptation and prediction inference speed in the target region is extremely fast, with the adaptation process taking only 10-30 seconds and the single-sample inference time being less than 1 second, meeting the timeliness requirements of practical engineering applications.
[0138] Example 2
[0139] The present invention also provides a small-sample geothermal resource prediction device, comprising:
[0140] The first processing module is used to generate multiple homogeneous synthesis tasks from a source dataset containing historical data of multiple geothermal regions by hierarchical category sampling and adding adaptive Gaussian noise.
[0141] The second processing module is used to automatically search for the highest priority hyperparameter of the Bayesian prior distribution in the homogeneous synthesis task by employing a Bayesian optimization algorithm based on Gaussian processes.
[0142] The third processing module is used to fix the highest priority hyperparameter on a small sample of data in a new target area, and only fine-tune the sharpening parameter, temperature parameter and upper limit of sample weight.
[0143] The fourth processing module is used to discretize the continuous geological features in the small sample data of the target area at three scales: coarse, medium and fine, and to construct a Bayesian prototype for each geothermal resource category at each scale based on the normal-inverse Wishart conjugate prior.
[0144] The fifth processing module is used to calculate the similarity between the sample to be predicted and historical samples at coarse, medium and fine scales using the Bayesian prototype. The coarse scale uses the discrete rank difference measure, the medium scale uses a hybrid measure of discrete rank and continuous feature posterior density, and the fine scale uses the posterior prediction density measure.
[0145] The sixth processing module is used to fuse the similarity of each scale according to a preset weight, and after sharpening and upper limit truncation, normalize to obtain the weight of each sample.
[0146] The seventh processing module is used to aggregate class weights based on sample weights and mix the Bayesian prediction distributions of each class into a mixed student model. The distribution is calculated, and the predicted mean and confidence intervals at the specified confidence level are output.
[0147] Example 3
[0148] The present invention also provides a small sample geothermal resource prediction system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a small sample geothermal resource prediction method when run by the processor.
[0149] Example 4
[0150] The present invention also provides a storage medium storing a computer program, which executes a small-sample geothermal resource prediction method when running.
[0151] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A small sample geothermal resource prediction method, characterized in that, The method comprises the following steps: Step S1, generating a plurality of homologous synthesis tasks from a source data set containing historical data of a plurality of geothermal regions by hierarchical category sampling and adding adaptive Gaussian noise; Step S2, automatically searching for optimal prior hyperparameters of a Bayesian prior distribution on the homologous synthesis tasks by using a Gaussian process-based Bayesian optimization algorithm; Step S3, fixing the optimal prior hyperparameters and only fine-tuning sharpening parameters, temperature parameters and upper limits of sample weights on small sample data of a new target region; Step S4, discretizing continuous geological features in the small sample data of the target region at three scales of coarse, medium and fine, and constructing Bayesian prototypes for each geothermal resource category at each scale based on a normal-inverse Wishart conjugate prior; Step S5, calculating the similarity of the to-be-predicted sample with the historical samples at the coarse, medium and fine scales by the Bayesian prototypes, wherein the coarse scale adopts a discrete level gap distance measurement, the medium scale adopts a hybrid measurement of discrete levels and continuous feature posterior density, and the fine scale adopts a posterior prediction density measurement; Step S6, fusing the similarity at each scale according to a preset weight, performing sharpening processing and upper limit truncation, and then normalizing to obtain each sample weight. Step S7, mix the Bayesian predictive distributions of each category into a mixture of Student's t-distributions according to the sample weights aggregated category weights, and output the predictive mean and the confidence interval at a specified confidence level. distribution, and output the predictive mean and the confidence interval at a specified confidence level.
2. A small sample geothermal resource prediction device, characterized by, The method comprises the following steps: A first processing module is configured to generate a plurality of homologous synthesis tasks from a source data set containing historical data of a plurality of geothermal regions by hierarchical category sampling and adding adaptive Gaussian noise; A second processing module is configured to automatically search for optimal prior hyperparameters of a Bayesian prior distribution on the homologous synthesis tasks by using a Gaussian process-based Bayesian optimization algorithm; A third processing module is configured to fix the optimal prior hyperparameters and only fine-tune sharpening parameters, temperature parameters and upper limits of sample weights on small sample data of a new target region; A fourth processing module is configured to discretize continuous geological features in the small sample data of the target region at three scales of coarse, medium and fine, and construct Bayesian prototypes for each geothermal resource category at each scale based on a normal-inverse Wishart conjugate prior; A fifth processing module is configured to calculate the similarity of the to-be-predicted sample with the historical samples at the coarse, medium and fine scales by the Bayesian prototypes, wherein the coarse scale adopts a discrete level gap distance measurement, the medium scale adopts a hybrid measurement of discrete levels and continuous feature posterior density, and the fine scale adopts a posterior prediction density measurement; A sixth processing module is configured to fuse the similarity at each scale according to a preset weight, perform sharpening processing and upper limit truncation, and then normalize to obtain each sample weight. a seventh processing module configured to mix the Bayesian prediction distribution of each category into a mixed Student distribution according to the sample weight aggregated category weight, and output the prediction mean and the confidence interval at a specified confidence level.
3. A small sample geothermal resource prediction system characterized by, The method comprises the following steps: A memory and a processor, wherein the memory stores a computer program which is run by the processor, and the computer program performs the small sample geothermal resource prediction method of claim 1 when being run by the processor.
4. A storage medium, characterized by The storage medium stores a computer program which performs the small sample geothermal resource prediction method of claim 1 when being run.
Citation Information
Patent Citations
Mineral resource classification prediction method and system based on multi-source small sample joint learning
CN116484295A
Small sample layered Bayesian estimation implementation method suitable for DSMC method flow field calculation
CN120951457A