A method and system for multi-resolution geological data conversion based on ensemble learning

By adopting a multi-resolution geological data conversion method based on ensemble learning, the spatial scale mismatch problem of geological data with different resolutions is solved, high-precision spatial prediction of target elements is achieved, reliable high-resolution geochemical maps are generated, and the efficiency of geological data utilization is improved. This method is applicable to fields such as mineral exploration and resource estimation.

CN121562863BActive Publication Date: 2026-04-17CHINA GEOLOGICAL SURVEY XIAN MINERAL RESOURCES SURVEY CENT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA GEOLOGICAL SURVEY XIAN MINERAL RESOURCES SURVEY CENT
Filing Date
2026-01-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to reasonably extrapolate unmeasured target elements in high-resolution spatial patterns without relying on additional field sampling and laboratory testing. This is especially true in western regions with lower levels of geological work, where traditional methods struggle to quickly fill data gaps through denser sampling. Existing remote sensing and geophysical methods cannot provide quantitative information on elemental concentration levels, and spatial scale mismatches and sampling density differences exist between geological data of different resolutions, leading to model learning biases and logical confusion.

Method used

A multi-resolution geological data conversion method based on ensemble learning is adopted. By acquiring low-resolution and high-resolution geological datasets, spatial grid aggregation processing is performed to train a Stacking ensemble regression model, generating a predictor variable matrix. Then, a fusion prediction model is generated using base learners and meta-learners, combined with spatial error correction, to achieve high-precision spatial prediction of target elements.

Benefits of technology

Reliable high-resolution geochemical maps can be generated at low cost without re-field sampling, improving the efficiency of geological data utilization, reducing redundant exploration investment, and making them suitable for mineral exploration, resource estimation, and environmental assessment, especially in areas with weak basic data but of great strategic significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562863B_ABST
    Figure CN121562863B_ABST
Patent Text Reader

Abstract

This application relates to a method and system for multi-resolution geological data conversion based on ensemble learning, belonging to the field of geological information processing technology. The conversion method includes: acquiring a low-resolution geological dataset containing target elements and a high-resolution geological dataset not containing target elements; performing spatial grid aggregation processing on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset; training a Stacking ensemble regression model with the predictor variable matrix as input and the target element values ​​in the low-resolution geological dataset as the output target; inputting the high-resolution geological dataset into the fusion prediction model to output preliminary predicted values ​​of the target elements; and performing spatial error correction on the preliminary predicted values ​​of the target elements based on the actual values ​​of the target elements in the low-resolution geological dataset to generate corrected high-resolution target element data. This application can effectively integrate multi-scale, multi-source heterogeneous geological data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of geological information processing technology, and in particular to a multi-resolution geological data conversion method and system based on ensemble learning. Background Technology

[0002] With the continuous advancement of my country's deep mineral exploration strategy, high-precision, comprehensive geochemical data has become a crucial foundation for supporting mineralization prediction and resource potential assessment. However, for a long time, geological surveys have generally faced a contradiction between data acquisition scale and testing indicators: while large-scale regional geochemical surveys can provide systematic data on multiple elements, their spatial resolution is low due to limitations in sampling density and analysis costs, making it difficult to meet the needs of detailed local exploration; while high-density local surveys, although possessing excellent spatial representation capabilities, often fail to comprehensively determine all key mineralization indicator elements due to funding, technology, or historical reasons, resulting in missing target element information in key areas. This data structural discontinuity severely restricts the accurate identification of geochemical anomalies and the in-depth analysis of their spatial distribution patterns.

[0003] Against this backdrop, how to make full use of existing geological data resources and achieve reasonable extrapolation of unmeasured target elements in high-resolution spatial patterns without relying on new field sampling and laboratory testing has become a key technical bottleneck restricting the upgrading of geological information services.

[0004] Especially in the vast western regions where geological work is relatively limited, traditional methods struggle to quickly fill data gaps through denser sampling, while existing indirect methods such as remote sensing and geophysics cannot directly provide quantitative information at the elemental concentration level. Although some studies in recent years have attempted to use ordinary interpolation methods or multiple regression models for elemental prediction, the highly nonlinear, multi-scale coupled, and complex geological processes-controlled nature of geochemical fields often make it difficult for single models to effectively capture deep correlations between variables, and they cannot simultaneously account for macroscopic statistical consistency and microscopic spatial heterogeneity. Furthermore, spatial scale mismatches, sampling density differences, and coordinate system biases between geological data of different resolutions further increase the technical difficulty of cross-scale modeling. When low-resolution and high-resolution geological data cannot be aligned in spatial units, directly establishing predictive relationships will lead to positional misalignments between training samples and labels, resulting in model learning biases and even logical confusion. A deeper problem is that, even if spatial matching is achieved, how to accurately transmit the full elemental spectrum information contained in low-resolution geological data, especially the enrichment trend of rare or key metal elements, while preserving the geomorphic response characteristics and local structural information contained in high-resolution geological data, remains a technical problem that urgently needs to be solved in current geological modeling. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method and system for multi-resolution geological data conversion based on ensemble learning.

[0006] Firstly, this application provides a multi-resolution geological data conversion method based on ensemble learning, employing the following technical solution:

[0007] Obtain a low-resolution geological dataset containing the target element, and a high-resolution geological dataset that does not contain the target element;

[0008] Spatial grid aggregation is performed on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset;

[0009] Using the predicted variable matrix as input and the target element values ​​in the low-resolution geological dataset as output, a Stacking ensemble regression model is trained; wherein, the first layer of base learners outputs the base learner prediction matrix, and the second layer of meta-learners generates a fusion prediction model based on the base learner prediction matrix.

[0010] The high-resolution geological dataset is input into the fusion prediction model, which outputs preliminary predicted values ​​of the target elements.

[0011] Based on the actual values ​​of the target elements in the low-resolution geological dataset, spatial error correction is performed on the preliminary predicted values ​​of the target elements to generate corrected high-resolution target element data.

[0012] By employing the aforementioned technical solution, target element information from low-resolution geological data and spatial details from high-resolution geological data are effectively integrated, enabling high-precision spatial prediction of target elements. Its practical significance lies in overcoming the information gaps caused by limitations in testing projects in traditional geological mapping. It generates reliable high-resolution geochemical maps at low cost without requiring re-sampling in the field, providing crucial data support for mineral exploration, resource estimation, and environmental assessment. This technical solution not only improves the utilization efficiency of existing geological data and reduces redundant exploration investment, but also provides an intelligent solution for constructing high-precision geochemical maps at the national scale, making it particularly suitable for strategically important regions such as western China where basic data is scarce.

[0013] Secondly, this application provides a multi-resolution geological data conversion method based on ensemble learning, employing the following technical solution:

[0014] Obtain a 3D geological body dataset;

[0015] The three-dimensional geological body dataset is layered vertically to generate multiple two-dimensional slice datasets;

[0016] Each two-dimensional slice dataset is input as a low-resolution geological dataset or a high-resolution geological dataset, and the multi-resolution geological data conversion method described in the first aspect is executed independently to obtain the corrected two-dimensional target element data.

[0017] All generated two-dimensional target element data are spatially reconstructed along the vertical direction to generate a three-dimensional target element data volume with uniform resolution.

[0018] By adopting the above technical solution, high-precision modeling and intelligent resolution upscaling of complex geological bodies under multi-scale conditions were achieved. The entire process fully utilizes the advantages of modern machine learning in nonlinear relationship mining, and incorporates the basic principles of geological sequence structure, vertical zonation, and spatial continuity, forming a dual support mechanism of "data-driven + knowledge-guided". Compared with traditional methods that rely on manual interpolation or empirical extrapolation, this technical solution improves the prediction accuracy and reliability of results for deep, unknown areas, and is particularly suitable for major engineering scenarios requiring detailed 3D characterization, such as mineral prospect prediction, groundwater resource assessment, and carbon sequestration site selection. The final output of a unified resolution 3D target element data volume not only possesses a complete spatial topological structure and physical consistency, but can also directly serve reserve calculation, risk assessment, and decision support systems, demonstrating the core value and broad prospects of artificial intelligence technology in promoting the digital transformation of geological science.

[0019] Thirdly, this application provides a multi-resolution geological data conversion system based on ensemble learning, employing the following technical solution:

[0020] The data acquisition module is used to acquire low-resolution geological datasets containing the target elements, as well as high-resolution geological datasets that do not contain the target elements.

[0021] The spatial aggregation and alignment module is used to perform spatial grid aggregation processing on the high-resolution geological dataset to generate a predictive variable matrix that is spatially aligned with the low-resolution geological dataset.

[0022] The model training and ensemble module is used to train a Stacking ensemble regression model with the predicted variable matrix as input and the target element values ​​in the low-resolution geological dataset as output targets; wherein, the first layer base learner group outputs the base learner prediction matrix, and the second layer meta learner generates a fusion prediction model based on the base learner prediction matrix.

[0023] The high-resolution prediction output module is used to input the high-resolution geological dataset into the fusion prediction model and output preliminary predicted values ​​of the target elements.

[0024] The spatial error correction module is used to perform spatial error correction on the preliminary predicted values ​​of the target elements based on the actual values ​​of the target elements in the low-resolution geological dataset, and generate corrected high-resolution target element data.

[0025] Fourthly, this application provides a multi-resolution geological data conversion system based on ensemble learning, employing the following technical solution:

[0026] The 3D data acquisition module is used to acquire 3D geological body datasets;

[0027] The slicing module is used to layer the three-dimensional geological body dataset along the vertical direction to generate multiple two-dimensional slice datasets.

[0028] The two-dimensional data correction module is used to take each two-dimensional slice dataset as input as a low-resolution geological dataset or a high-resolution geological dataset, and independently execute the multi-resolution geological data conversion method as described in the first aspect to obtain the corrected two-dimensional target element data.

[0029] The spatial reconstruction module is used to spatially reconstruct all generated two-dimensional target element data along the vertical direction, generating a three-dimensional target element data volume with uniform resolution.

[0030] Fifthly, this application provides a computer-readable storage medium, which adopts the following technical solution:

[0031] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the first process of a multi-resolution geological data conversion method based on ensemble learning, which is one embodiment of this application.

[0033] Figure 2 This is a comparison diagram of sampling density at different scales in one embodiment of this application.

[0034] Figure 3 This is a data structure diagram of a large-scale dataset and a small-scale dataset according to one embodiment of this application.

[0035] Figure 4 This is a technical roadmap for converting small-scale data into large-scale data, according to one embodiment of this application.

[0036] Figure 5 This is a schematic diagram of the second process of a multi-resolution geological data conversion method based on ensemble learning, which is one embodiment of this application.

[0037] Figure 6This is a schematic diagram of the third process of a multi-resolution geological data conversion method based on ensemble learning, according to one embodiment of this application.

[0038] Figure 7 This is a schematic diagram of the fourth process of a multi-resolution geological data conversion method based on ensemble learning, one embodiment of this application.

[0039] Figure 8 This is a schematic diagram of the fifth process of a multi-resolution geological data conversion method based on ensemble learning, according to one embodiment of this application.

[0040] Figure 9 This is a schematic diagram of the sixth process of a multi-resolution geological data conversion method based on ensemble learning, according to one embodiment of this application.

[0041] Figure 10 This is a schematic diagram of the seventh process of a multi-resolution geological data conversion method based on ensemble learning, according to one embodiment of this application. Detailed Implementation

[0042] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-10 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0043] This application discloses a multi-resolution geological data conversion method based on ensemble learning.

[0044] Reference Figure 1 A multi-resolution geological data conversion method based on ensemble learning, comprising:

[0045] Step S101: Obtain a low-resolution geological dataset containing the target element and a high-resolution geological dataset not containing the target element.

[0046] Low-resolution geological datasets typically refer to large-scale measurement results with low sampling density and wide spatial coverage, such as 1:200,000 scale regional geochemical survey data. Although such data have fewer sample points per unit area, resulting in blurred spatial details, they can often comprehensively detect a variety of trace elements due to relatively controllable analysis costs, thus preserving the concentration values ​​of key target elements (such as mineralization indicator elements like gold, copper, and lead).

[0047] High-resolution geological datasets represent more precise local survey data, such as water system sediment and soil measurements at a scale of 1:50,000 or higher. They are densely sampled and spatially accurate, and can finely characterize the spatial variation structure of surface materials. However, due to economic and technical limitations, not all elements are measured, and certain specific target elements may be missing.

[0048] Understandably, the two types of data are complementary: low-resolution geological data provides "what," while high-resolution geological data reveals "where." By combining the two, it is possible to reasonably reconstruct the distribution of target elements in high-resolution space without resampling in the field, thereby overcoming the information gap problem caused by the limitation of test items in traditional geological mapping.

[0049] Reference Figure 2 This paper compares the sampling density of stream sediment data at different scales. Within a 4 square kilometer area, one sample was collected at a scale of 1:200,000, four samples at 1:100,000, and 16 samples at 1:50,000. Clearly, smaller scale data (e.g., 1:200,000) typically include more elemental analyses but have lower precision; larger scale data (e.g., 1:50,000) have fewer elemental analyses but higher precision, better reflecting the spatial distribution of elements. This comparison highlights the necessity of data transformation.

[0050] Reference Figure 3 This illustrates the differences and correlations in data structure between the input and target data. The right side of the figure shows a high-precision, large-scale dataset X, which contains spatial distribution information for multiple elements (a, b, c, ..., n). The dense data points provide a detailed characterization of geochemical features. However, this dataset lacks the specific target element y that we want to predict. The left side shows a small-scale dataset Y, which contains the information for this target element Y. However, due to its small scale, its spatial resolution is low, and each data point represents an average value over a larger range, failing to capture detail. The core idea of ​​this application is to utilize a machine learning model to establish a complex mapping relationship between the known high-resolution multi-element data (X) and the low-resolution target element data (Y), thereby predicting the missing high-resolution target element y.

[0051] In this embodiment of the application, the large-scale dataset X contains elements a, b, cn, etc., but lacks element y; while the small-scale data corresponds to element Y.

[0052] Step S102: Perform spatial grid aggregation processing on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset.

[0053] This step is essentially a spatial scale matching operation, designed to establish comparability and modeling feasibility between two geological data sets with different resolutions. Because high- and low-resolution geological data differ significantly in their original sampling layout, coordinate systems, and spatial units, direct modeling would result in a lack of correspondence between input and output.

[0054] Therefore, a geographic grid of low-resolution geological data (usually a regular rectangular grid, such as 4 square kilometers per grid) is used as a reference frame. All sampling points falling within the same grid in the high-resolution geological data are then aggregated and calculated. The most common method is to take the arithmetic mean, that is, to calculate the mean of multiple geochemical indicators (such as non-target elements like Fe, Mn, SiO2, and K2O) at all high-resolution points within a grid, forming a unified vector representation. The resulting "predictor variable matrix" is actually a dimensionality-reduced spatially aligned dataset. Each row corresponds to a low-resolution grid cell, and each column represents the aggregated value of a geochemical variable. It retains the local geological background characteristics reflected by the high-resolution geological data while achieving strict spatial registration with the actual observed values ​​of the target elements. This aggregation is not only a geometric resampling but also a process of information abstraction: it extracts stable regional trend components from the high-resolution geological data, weakens the influence of random noise, and provides physically meaningful and statistically robust input features for subsequent machine learning models.

[0055] Step S103: Train the Stacking ensemble regression model with the prediction variable matrix as input and the target element values ​​in the low-resolution geological dataset as output targets; wherein, the first layer base learner group outputs the base learner prediction matrix, and the second layer meta learner generates the fusion prediction model based on the base learner prediction matrix.

[0056] Specifically, Stacking, as an advanced ensemble learning strategy, is designed based on the machine learning philosophy that "collective intelligence is superior to individual judgment" and aims to capture complex nonlinear mapping relationships through multi-level model collaboration.

[0057] Specifically, the first layer consists of multiple heterogeneous base learners, including but not limited to classic machine learning models such as random forests, support vector machines, and gradient boosting trees. These models each possess different inductive preferences: random forests excel at handling high-dimensional sparse features and resist overfitting; support vector machines can effectively delineate complex boundaries even with small sample sizes; and gradient boosting trees achieve strong fitting capabilities through iterative optimization of residuals. Each base learner independently learns a mapping function from the predicted variable to the target element value on the training set and outputs its preliminary prediction result for the target. The set of predictions from all base learners constitutes the "base learner prediction matrix." This matrix no longer uses the original geochemical variables as its dimension, but rather the cognitive perspective of each model as its new feature space, essentially transforming the original geological information into a set of "model-level semantic features."

[0058] The second layer introduces a meta-learner (usually linear regression), whose task is not to directly fit the original input, but to learn how to optimally weight and combine the predictions from different models to generate the final "fusion prediction model." This model not only inherits the advantages of each base model, but also automatically identifies which models are more reliable under what conditions, and then dynamically adjusts the weight allocation, greatly improving the overall prediction accuracy and generalization performance. More importantly, due to the use of cross-validation for parameter tuning, the entire modeling process avoids the risk of overfitting, ensuring that the learned patterns have real extrapolation capabilities.

[0059] Step S104: Input the high-resolution geological dataset into the fusion prediction model and output the preliminary predicted values ​​of the target elements;

[0060] In this process, the original high-resolution geological dataset is input into the fusion prediction model that has been trained, and the model outputs preliminary predicted values ​​of the target elements, marking the transition of the model from the "learning stage" to the "application stage". At this point, the model no longer operates on the aggregated coarse-grained data, but directly deals with the unprocessed high-density sampling point sequence.

[0061] For each high-resolution sampling location, the system extracts surrounding geological variables to form an input vector, which is then fed into the fusion prediction model for forward inference to obtain the estimated concentration of the target element at that location. Since the model has already learned deep correlation patterns between the target element and various associated elements (e.g., enrichment of a certain metal is often accompanied by specific silicate mineral assemblages), even if the original high-resolution geological data itself did not determine the element, it can make reasonable inferences based on its geological environment characteristics. The resulting "preliminary predicted value of the target element" is a new spatially continuous and detailed data layer, theoretically possessing the same spatial resolution as the high-resolution geological data, thus initially realizing the knowledge transfer from low-resolution information to high-resolution representation.

[0062] However, due to the information loss inherent in the aggregated data used during training, and the potential for systematic biases in the model (such as an overall tendency to overestimate or underestimate), the prediction results cannot be directly used for quantitative interpretation and must be further corrected.

[0063] Step S105: Based on the actual values ​​of the target elements in the low-resolution geological dataset, perform spatial error correction on the preliminary predicted values ​​of the target elements to generate corrected high-resolution target element data.

[0064] This step is not a simple numerical scaling, but a spatial consistency adjustment based on the principle of global mass conservation. The basic logic is that although high-resolution predictions provide rich spatial details, on a large-scale average level, their total content should be consistent with the known total content measured at low resolution; otherwise, it would violate geological facts.

[0065] Therefore, by calculating the regional error correction factor λ, which is the ratio of the measured total amount of low-resolution target elements to the preliminary predicted total amount, the overall degree of deviation can be quantified. Then, this factor is combined with the spatial weight of each grid (considering differences in grid area or sampling density) to correct the predicted values ​​grid by grid. This ensures that the final output of "high-resolution target element data" not only exhibits reasonable spatial fluctuations at the microscopic level but also faithfully reproduces the true geochemical background at the macroscopic statistical level. This correction method balances local morphological fidelity with overall numerical consistency, effectively preventing spurious anomalies or signal attenuation caused by model extrapolation. The generated data can be used for both mineralization prospect delineation and quantitative analysis tasks such as resource estimation.

[0066] The above implementation effectively integrates target element information from low-resolution geological data with spatial details from high-resolution geological data, achieving high-precision spatial prediction of target elements. Its practical significance lies in overcoming the information gaps caused by limitations in testing projects in traditional geological mapping. It generates reliable high-resolution geochemical maps at low cost without requiring re-sampling in the field, providing crucial data support for mineral exploration, resource estimation, and environmental assessment. This technical solution not only improves the utilization efficiency of existing geological data and reduces redundant exploration investment, but also provides an intelligent solution for constructing high-precision geochemical maps at the national scale, making it particularly suitable for strategically important regions such as western China where basic data is scarce.

[0067] Reference Figure 4 This document details the complete machine learning workflow for transforming small-scale data into large-scale data, presented in the form of a technology roadmap. The workflow primarily comprises four key steps:

[0068] Step 1: Data preparation and synthesis. By spatial averaging the high-precision, large-scale grid data, its resolution is matched with that of the low-precision, small-scale data, thereby constructing the predictor variable set X and the target variable set Y for training the model.

[0069] Step 2: Model Building and Training. This is the most complex step, employing the advanced Stacking ensemble learning framework. This framework consists of two layers: the first layer is the "base learner layer," which uses multiple machine learning algorithms in parallel (such as logistic regression, support vector machines, random forests, etc.). Each algorithm undergoes rigorous hyperparameter optimization to learn different patterns from the data. The second layer is the "meta learner layer" (usually using linear regression). Its role is not to directly learn from the raw data, but rather to learn how to most effectively weight and combine the predictions of all models in the first layer to obtain a stronger and more stable final prediction model.

[0070] Step 3: Apply the trained ensemble model to predict the original high-resolution large-scale data (X) to obtain a preliminary distribution map of the high-resolution target element Y(x).

[0071] Step 4: Result correction. A redistribution formula is used to ensure that the sum of all predicted grid point values ​​is consistent with the true total value of the original small-scale data, thereby ensuring the quality conservation of data conversion and making the prediction results more consistent with the actual geological situation.

[0072] Reference Figure 5 As one implementation of step S102, the step of performing spatial grid aggregation processing on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset includes:

[0073] Step S201: Define the target grid cell partitioning rules based on the spatial grid structure of the low-resolution geological dataset;

[0074] Low-resolution geological data is typically organized in the form of regular geographic grids, such as the 4-square-kilometer standard map sheet formed based on a 1:200,000 scale regional geochemical survey. These grids not only have fixed latitude and longitude boundaries or projected coordinate ranges, but also represent the basic units of the original sampling design. Using these existing grid structures directly as the basis for dividing "target grid units" means that all information from the high-resolution geological data will be forcibly mapped into a spatial container that perfectly matches the low-resolution geological data, thereby establishing a strict "one-to-one" spatial correspondence.

[0075] Understandably, this step avoids registration errors and topological inconsistencies caused by custom grids, ensuring that data from different sources can be compared and modeled within the same geographic framework. More importantly, since the target element values ​​(such as the content of a certain metal) in low-resolution geological data are themselves published in units of that grid, using it as a benchmark to construct the predictor variable matrix naturally satisfies the basic requirement of "input and output in the same domain" in supervised learning. This allows the model training to be based on truly observable geographic units, improving the interpretability and reliability of the modeling results.

[0076] Step S202: Based on the target grid cell partitioning rules, map the data of each sampling point in the high-resolution geological dataset to the corresponding target grid cell;

[0077] Mapping each sampling point in a high-resolution geological dataset to its corresponding target grid cell is a precise allocation process based on spatial location matching. High-resolution geological data often originates from denser field sampling, such as 1:50,000 stream sediment measurements, which may contain dozens or even hundreds of sampling points in the same area. These points are spatially distributed as discrete points and are not laid out according to a fixed grid.

[0078] To ensure that the information effectively serves modeling tasks using coarse grids, it is necessary to determine the target grid to which each high-resolution sampling point belongs through coordinate analysis. This process relies on spatial indexing and point-to-polygon discrimination algorithms in Geographic Information Systems (GIS), specifically determining the grid affiliation by checking whether the geographic coordinates of a point fall within a certain polygon (target grid).

[0079] This mapping is not only a geometric operation but also a prerequisite for information aggregation. It determines which high-resolution samples will participate in the feature representation of the same coarse grid. Although high-resolution geological data do not determine the target elements, they contain rich associated components (such as Fe, Al, SiO2, etc.), mineral assemblages, or geophysical response signals. These variables often have a geological genetic correlation with the target elements in a local range. Therefore, classifying them by grid is essentially extracting a "local geological background fingerprint," providing a data foundation for the subsequent inversion of unknown elements from background information.

[0080] Step S203: Perform aggregation calculation on all sampled point element values ​​within each target grid cell to generate aggregated feature values ​​for each target grid cell;

[0081] The core step in achieving information concentration and noise suppression is to perform aggregation calculations on the element values ​​of all sampled points within each target grid cell to generate the aggregated feature value of that grid cell. Since a single target grid cell may contain multiple high-resolution sampled points, directly retaining the original point data would lead to an expansion of the input dimension and make the model difficult to process; while randomly selecting a point would result in the loss of a large amount of spatial variation information.

[0082] Therefore, using aggregation calculations (such as arithmetic mean) is a reasonable and robust choice. By averaging the element values ​​of all sampling points within the same grid, local anomalies and measurement errors can be smoothed out while preserving the overall trend, resulting in a comprehensive index representing the typical geochemical characteristics of the region. This aggregation method is essentially a "spatial weighted summary" of high-resolution geological data, ensuring that the final feature vector of each grid reflects the average attribute state of the geological bodies within it.

[0083] The specific aggregation calculation formula is as follows:

[0084] ;

[0085] In the above formula, N represents the aggregated feature value of the i-th grid cell. i E represents the number of sampling points within this unit. i,k Let be the element value of the kth sampling point.

[0086] It is worth noting that although the averaging method is the most commonly used aggregation strategy, in practical applications, the median, geometric mean, or weighted average can be selected according to the geological background to adapt to skewed distributions or situations controlled by specific tectonic factors, demonstrating the flexibility and scalability of this method in engineering implementation.

[0087] Step S204: Integrate the aggregated feature values ​​of all target grid cells to form a predictor variable matrix.

[0088] The predictor variable matrix is ​​a structured two-dimensional array, where each row corresponds to a target grid cell defined by low-resolution geological data, and each column represents the aggregated value of a geochemical element or other geological parameter. For example, if there are M target grids and P analysis elements, an M×P dimension matrix is ​​generated, which can be used as the input feature set for the supervised learning model.

[0089] In this embodiment, the matrix is ​​not only a set of variables in a mathematical sense, but also a spatial encoding carrier of geological knowledge: it transforms the originally scattered, unstructured high-density observation data into a set of predictive factors that are strictly aligned with the target response variables (i.e., the measured values ​​of target elements in low-resolution geological data), thereby establishing a learnable mapping path between "known background and unknown target". It is this structured representation that enables subsequent ensemble learning models to effectively uncover the complex nonlinear relationships hidden between multidimensional geological variables and to be used for refined extrapolation in high-resolution space.

[0090] The above implementation not only solves the problem of scale and structure mismatch in multi-source heterogeneous geological data, but also achieves the goal of extracting useful information from high-precision spatial sampling and serving low-precision target prediction without introducing additional assumptions, significantly enhancing the model's generalization ability and the reliability of the results.

[0091] Reference Figure 6 As one implementation of step S103, the steps of training the Stacking ensemble regression model, with the predictor variable matrix as input and the target element values ​​in the low-resolution geological dataset as the output target, include:

[0092] Step S301: Construct a first-layer base learner group containing multiple base learners, and train each base learner separately using the prediction variable matrix and the target element values ​​in the low-resolution geological dataset.

[0093] Base learners refer to a set of basic machine learning models with different structures and inductive preferences, such as random forests, gradient boosting trees, and support vector machines. These models exhibit different advantages when processing geological data: random forests excel at capturing the interactions between high-dimensional features and are robust to outliers; gradient boosting trees gradually approximate the true function by iteratively fitting residuals, making them particularly suitable for data with strong nonlinear trends; while support vector machines can effectively delineate complex decision boundaries even with small sample conditions, making them suitable for scenarios in regional surveys with limited sampling points but high variable dimensionality.

[0094] Specifically, grouping these models side-by-side into a first-layer model group means that the system no longer relies on the assumptions of a single model, but allows the same set of input-output relationships to be observed simultaneously from multiple mathematical perspectives. Each base learner independently learns the mapping between the predictor variable (i.e., the aggregated geochemical background information) and the target element concentration on the same training set, but due to their different internal mechanisms, their understanding of the data, the key areas they focus on, and the characteristics of their error distribution also differ.

[0095] Step S302: Input the prediction variable matrix into each base learner to generate the corresponding base learner prediction results;

[0096] This step marks the transition of the model from the "learning phase" to the "inference phase." Although the input data from the training phase (the training set used to build the meta-learner) is still used, the goal at this point is no longer parameter updating, but rather extracting the estimated output of each base model for the target variable. Since each base learner has completed its optimal fit within its own structure, the predicted values ​​it produces reflect the model's best judgment on the spatial distribution of the target element under the current geological environment.

[0097] It's important to note that these predictions do not aim for complete consistency; rather, a certain degree of divergence is retained. This is because moderate inter-model variation helps reveal areas of uncertainty in the data. For example, in grid cells with complex geological structures or ambiguous mineralization conditions, some models may predict higher values, while others tend to be more conservative. This "disagreement" itself is valuable information and will be appropriately weighted during subsequent fusion. Therefore, this step is not only a necessary part of the technical process but also a crucial transition for achieving knowledge transfer and semantic enhancement between models.

[0098] Step S303: Integrate the prediction results of all base learners to form a base learner prediction result matrix;

[0099] This step is crucial for transitioning from "individual wisdom" to "collective cognition." The prediction matrix uses each low-resolution grid cell as a row index and the prediction output of each base learner as a column feature, forming a structured two-dimensional array. In this new feature space, the original geochemical variables have been transformed into a set of "model-level response signals," meaning each grid is no longer described by its physical composition but by its comprehensive performance across multiple algorithmic perspectives. This transformation is akin to "translating" raw observational data into a higher-level semantic expression, allowing complex geological processes that were previously difficult to model directly to be indirectly characterized through model consensus.

[0100] More importantly, this matrix, as input to the meta-learner, actually constitutes a completely new training set, where each row represents the "model fingerprint" of a geographic unit, and each column reflects the cognitive tendency of a certain type of algorithm. This design allows the meta-learner to no longer directly face the original geochemical measurements with high noise levels, but instead to relearn based on the collective judgments of multiple expert systems, thereby significantly improving the stability and generalization ability of the final model.

[0101] Step S304: Using the prediction result matrix of the base learner as the input feature and the target element value as the output target, train the meta learner to generate the fusion prediction model.

[0102] In this process, the predicted result matrix is ​​used as the input feature, and the actual values ​​of the low-resolution target elements are used as the output target. A meta-learner is trained to generate a fusion prediction model, thus completing the closed-loop construction of the entire Stacking framework. The meta-learner plays the role of a "coordinator" or "referee" here. Its task is not to rediscover the underlying rules, but to learn how to optimally combine the opinions from different base learners.

[0103] In the embodiments of this application, linear regression is used as a meta-learner, which is a preferred strategy that balances interpretability and computational efficiency: by assigning appropriate weight coefficients to the prediction results of each base model, it forms a final prediction in the form of a weighted average, which can both retain the advantages of each model and suppress the extreme bias of individual models.

[0104] More importantly, the training process still relies on supervised learning based on real-world observed target values, ensuring that the fused prediction results faithfully reproduce the statistical characteristics of the measured data. Since the meta-learner is trained using a cross-validation mechanism, its weight allocation exhibits excellent extrapolation performance, maintaining consistent performance across new regions or unknown samples. The resulting "fusion prediction model" is no longer a simple averager, but an adaptive intelligent ensemble capable of dynamically adjusting the influence of each sub-model according to different geological backgrounds, achieving true "adaptability to changing circumstances."

[0105] In the above implementation, a two-layer cascaded Stacking ensemble regression model is constructed, enabling efficient mining and reliable inference of hidden patterns in multi-source heterogeneous geological data. From the parallel modeling of the first-layer multi-base learners to the intelligent fusion of the second-layer meta-learners, the entire training process fully utilizes the unique advantages of various machine learning algorithms while avoiding the limitations of a single model through a structured ensemble mechanism. This technical solution not only improves prediction accuracy and stability but, more importantly, enhances the model's adaptability and interpretability in complex geological environments, providing solid technical support for the intelligent generation of high-resolution geological data.

[0106] Overall, this model training strategy fully embodies the advanced concept of "data-driven approach, model collaboration as a means, and geological authenticity as the goal," significantly improving the automation level and scientific value of traditional geological data processing, and has broad application prospects and promotional significance.

[0107] Reference Figure 7 As one implementation of step S104, the step of inputting a high-resolution geological dataset into the fusion prediction model and outputting preliminary predicted values ​​of the target elements includes:

[0108] Step S401: Extract the spatial unit structure of the high-resolution geological dataset;

[0109] Spatial unit structure refers to the basic geographical unit of division upon which high-resolution geological data is acquired and organized. This typically manifests as regular or irregular grids, sets of sampling points, or watershed units. For example, in a 1:50,000 scale stream sediment survey, each sample point constitutes an independent spatial unit; while in soil geochemical surveys, regular grids may be laid out in units of square kilometers. These spatial units are not only the smallest carriers of data storage but also the spatial benchmarks for subsequent modeling and interpretation.

[0110] Understandably, the purpose of extracting this structure is to clarify the granularity of the prediction operation—that is, each target to be predicted will be bound to a specific and identifiable geographical location. This approach avoids the risks of data misalignment, overlap, or omission, ensuring the feasibility of direct visualization and further analysis of the prediction results in a Geographic Information System (GIS). More importantly, because this structure originates directly from the original high-resolution geological data itself, its spatial distribution density is much higher than that of low-resolution geological data. Therefore, it can provide rich levels of detail for the final generated prediction maps, thus truly achieving information enhancement in the sense of "upsizing."

[0111] Step S402: Input the element data of each spatial unit independently into the fusion prediction model;

[0112] The process of independently inputting elemental data from each spatial unit into the fusion prediction model marks the formal entry of the model into the inference and execution phase. Elemental data specifically refers to all measured geochemical variables other than the target element, such as associated components like Fe, Mn, SiO2, K2O, Th, and U. These together constitute a key feature vector describing the local geological background. Although these variables did not include the target element (such as Au, Cu, Pb) in the original survey, their combination patterns are often closely related to geological processes such as mineralization, weathering degree, and parent rock type, thus containing clues that can be used to indirectly infer the enrichment state of the target element.

[0113] In this embodiment, by feeding the complete elemental spectrum of each spatial unit into the fusion prediction model, the model essentially determines the most likely concentration level of the target element under the current geochemical environment based on the complex nonlinear mapping relationships learned during the training phase. It should be noted that independent input means that the predictions of each unit do not affect each other, which conforms to the local stationarity assumption in geoscientific spatial autocorrelation modeling, and also facilitates parallel computation to improve efficiency.

[0114] Furthermore, since the fusion prediction model is essentially built on the Stacking framework, it integrates the cognitive results of multiple base learners and performs optimal weighted fusion by the meta-learner. Therefore, this step is not a simple function substitution, but a highly intelligent knowledge transfer process: it accurately projects the regular cognition established at the low-resolution scale to the specific location in the high-resolution space, realizing cross-scale knowledge generalization.

[0115] Step S403: Based on the regression calculation of the fusion prediction model, generate the target element prediction value for each spatial unit;

[0116] The technical essence of this step is to perform an end-to-end forward inference process, with the most crucial part being led by the meta-learner. When the element data of a spatial unit is fed into the model, it first triggers the independent prediction behavior of each base learner in the first layer, generating a set of preliminary estimates of the target element. These intermediate results are then aggregated into a new feature vector, which serves as the input to the meta-learner. Based on the weight configuration learned during the training phase, the meta-learner weights and combines these heterogeneous prediction results, ultimately outputting a comprehensive regression estimate.

[0117] The advantage of this mechanism is that it not only absorbs the sensitivity of random forests to nonlinear interactions and the advantages of support vector machines in boundary identification, but also utilizes gradient boosting trees to capture residuals more precisely, thus forming a more robust and accurate predictive output than any single model.

[0118] More importantly, because the meta-learner is trained using real target values ​​under strict supervision, its predictions exhibit good consistency in statistical trends and can approximate the actual geochemical distribution patterns at the macroscopic level. Therefore, the predicted target element values ​​for each spatial unit are not based on guesswork, but rather on scientific inferences validated by a large amount of historical data, possessing strong geological credibility.

[0119] Step S404: Integrate the target element prediction values ​​of all spatial units to form a preliminary target element prediction value set.

[0120] This step completes the closed-loop construction from "individual prediction" to "global representation." By reorganizing all predicted values ​​into a unified dataset according to the geographical order or coordinate index of the original spatial units, it ensures that the output results are completely consistent with the original high-resolution geological data in terms of spatial topology. It preserves the original sampling density and distribution pattern while adding previously missing target element information. This "attribute injection without changing the structure" approach allows the newly generated data to be directly embedded into existing geological databases, used in 3D modeling, or for anomaly delineation without additional format conversion or spatial registration.

[0121] Meanwhile, this set of predicted values ​​also provides the necessary input basis for the next step of spatial error correction. Only with continuous prediction across the entire domain can the systematic deviation between the predicted values ​​and the low-resolution measured values ​​be effectively evaluated, and global scaling adjustments or local corrections be implemented accordingly, so as to ensure that the final results have both high-resolution details and macroscopic realism.

[0122] The above implementation method achieves a substantial transformation from a trained fusion prediction model to a high-resolution target element distribution map. This process not only fully utilizes the knowledge accumulated in the previous modeling but also ensures a comprehensive improvement in spatial accuracy, geological rationality, and engineering usability of the prediction results through a rigorous unit-level processing strategy. This prediction execution flow enables the high-fidelity reconstruction of target element information that originally existed only in coarse-scale observations at fine-scale spaces, expanding the application boundaries of traditional geological survey data and providing powerful intelligent support for mineral resource potential assessment, environmental geochemical monitoring, and digital land construction.

[0123] Reference Figure 8 As one implementation of step S105, the step of performing spatial error correction on the preliminary predicted values ​​of the target elements based on the actual values ​​of the target elements in the low-resolution geological dataset to generate corrected high-resolution target element data includes:

[0124] Step S501: Calculate the regional scale error correction coefficient based on the actual values ​​of the target elements in the low-resolution geological dataset;

[0125] The core objective of this step is to identify and quantify systematic biases in model predictions. Regional scale refers to the overall extent based on geographic units defined by low-resolution geological data, such as a 1:200,000 map sheet or an administrative division.

[0126] Within this range, the sum of actual measurements of target elements in low-resolution geological data (i.e., the sum of all grid target values ​​within the region) is used as the true reference value. Simultaneously, the preliminary predictions generated by the high-resolution model are integrated and summed within the same geographic boundary to obtain the total prediction. The ratio between the two constitutes the error correction coefficient, which essentially reflects the overall amplification or reduction of the current prediction system relative to the actual situation. If this coefficient is greater than 1, it indicates that the prediction is too low and needs to be adjusted upwards; if it is less than 1, it indicates that the prediction is too high and should be compressed.

[0127] Understandably, this total-scale matching-based correction method differs from point-by-point correction. It focuses more on maintaining geological rationality on a large scale and avoiding global imbalance caused by local overfitting. Especially in applications such as mineral prediction, the accuracy of total resources is far more important than the location of individual anomalies. Therefore, the introduction of this coefficient effectively applies a "geological credibility anchor" to the model output, ensuring its consistency with existing survey results at the macro level.

[0128] Step S502: Determine the area weighting factor of each spatial unit based on the spatial grid distribution characteristics of the high-resolution geological dataset.

[0129] This step aims to achieve differentiated allocation of error correction in a fine-grained spatial manner. High-resolution geological data typically employs regular or irregular grid layouts, and the geographical area represented by each sampling cell may vary, especially in areas with projection transformations or complex topography. Simply applying the same correction coefficient uniformly to all cells would lead to larger cells being underestimated in their contribution, while smaller cells would be over-adjusted, thus distorting the true enrichment pattern of elements.

[0130] Therefore, an area weighting factor is introduced, which essentially measures the "spatial influence" of each spatial unit within its region. The larger the area, the higher its original contribution to the overall regional calculation, and thus it should bear a larger share of adjustment during error redistribution. This weighting factor is directly derived from the geometric attributes of each unit, such as its planar projected area or ellipsoidal surface area. It requires no manually set thresholds or classification rules, possessing complete calculability and objectivity. Through this mechanism, the system can rationally distribute the correction pressure to each sub-unit according to its spatial proportion while maintaining the total amount constant. This ensures that the final result conforms to the overall constraints while respecting the heterogeneous characteristics of the original spatial structure.

[0131] Step S503: Apply the error correction coefficient and area weighting factor to the preliminary predicted value of the target element to generate the spatial redistribution result;

[0132] This step is not a simple linear scaling, but rather a dual adjustment mechanism: on the one hand, it uses the error correction coefficient to guide the direction of the total prediction, and on the other hand, it uses the area weight factor to achieve spatial sensitivity regulation.

[0133] Specifically, the initial predicted value of each high-resolution spatial unit is multiplied by the combined effect of these two factors to form a new adjusted value. The significance of this joint modulation lies in its avoidance of a crude, one-size-fits-all correction, instead employing a gradual optimization path that aligns with geographical reality. For example, in a large area of ​​low background at the edge of a metallogenic belt, even if the predicted value is slightly higher, the adjustment will be relatively mild due to the large area weight; while high anomalies in narrow valleys, although small in area, can still maintain their relative intensity through a local fidelity preservation mechanism due to their crucial role in metallogenic information extraction. The entire process embodies the design concept of "macro-control, micro-adaptation," ensuring that the corrected data not only numerically matches the measured totals but also spatially closely approximates the true logic of geological evolution.

[0134] Step S504: Integrate the spatial redistribution results of all spatial units and output the corrected high-resolution target element data.

[0135] By reorganizing all adjusted predicted values ​​according to the spatial index order of the original high-resolution geological data, the output results are ensured to be completely consistent with the input data in terms of format, coordinate system, and resolution, without the need for additional interpolation or resampling. More importantly, while retaining the original sampling density and spatial details, this dataset has been given a new numerical system that conforms to the actual measured totals in the region, truly achieving the dual goals of "high accuracy + high fidelity". This type of data can be directly used for GIS platform visualization, 3D geological modeling, anomaly delineation, and resource potential evaluation, significantly improving the engineering transformation capability of traditional machine learning prediction results.

[0136] The above implementation not only effectively compensates for the shortcomings of purely data-driven models in terms of physical consistency, but also enhances the credibility and authority of the prediction results by introducing interpretable correction logic. This error correction strategy successfully achieves a deep integration of "model intelligence" and "geological laws," providing solid technical support for constructing a high-precision geochemical map at the national scale, multiple levels, and in an integrated manner.

[0137] Reference Figure 9 As a further implementation of the multi-resolution geological data conversion method, the conversion method also includes:

[0138] Step S601: Obtain a high-resolution geological dataset containing the target elements;

[0139] Among them, target elements refer to chemical elements or mineral components that need to be focused on in specific application scenarios. For example, in mineral exploration, they may be mineralization indicator elements such as Au, Cu, Pb, and Zn, while in environmental assessment, they may be heavy metal pollutants such as As, Cd, and Hg.

[0140] Step S602: Based on the preset low-resolution target grid structure, perform spatial grid cell division on the high-resolution geological dataset;

[0141] The pre-defined low-resolution target grid structure is often derived from existing authoritative geological maps or national / industry standard databases, such as the regular grid layout under a fixed projection coordinate system used in the National Geochemical Baseline Map. Accurately matching high-resolution geological data points (such as the latitude and longitude and elemental content of each sampling point) to this pre-defined grid system requires spatial overlay analysis techniques. Typical methods include point-in-polygon attribution, raster resampling, or centroid matching.

[0142] For example, when a soil sample with a 10-meter resolution falls within a 1-kilometer grid, the system automatically assigns it to that grid and marks it as an object to be aggregated. This process must strictly adhere to the principle of spatial topological consistency to avoid omissions, duplications, or misalignments. Especially in trans-zonal projections or edge areas, the distortion caused by coordinate transformations must be considered, and geographic registration corrections should be introduced when necessary. Through this step, the originally scattered and disordered high-density observation points are organized into coarse-grained spatial containers, forming a data mapping relationship of "fine-grained input—coarse-grained loading," preparing for the next step of statistical compression.

[0143] Step S603: Aggregate the element values ​​within each spatial grid cell after division to generate the aggregated value for each spatial grid cell;

[0144] This step is essential for information condensation and scale transition. Because geological elements often exhibit non-uniform spatial distribution, extreme outliers (such as mineralization centers) may exist in certain areas. Simply using an arithmetic average could distort the overall trend. Therefore, aggregation strategies need to be differentiated based on element type and its statistical distribution characteristics.

[0145] Specifically, for background elements that exhibit an approximately normal distribution (such as major elements like SiO2 and Al2O3), the arithmetic mean can effectively reflect the overall level within the region; while for trace elements that exhibit a skewed distribution or significant enrichment (such as Au and Ag), the median, geometric mean, or truncated mean is more suitable to suppress outlier interference; for categorical variables (such as rock type and mineral assemblage), the mode or category frequency distribution can be used to express them.

[0146] Furthermore, weighting factors can be set by incorporating geological knowledge, such as inverse distance weighting (IDW) based on the distance of sampling points from the grid center, giving samples closer to the center a higher influence. This intelligent aggregation method retains key geological information while filtering out local noise, achieving a transformation from "detailed but redundant" to "simple, clear, and representative".

[0147] Step S604: Integrate the aggregated values ​​of all spatial grid cells to generate a synthetic low-resolution geological dataset.

[0148] The generated "synthetic low-resolution geological dataset" should meet the various specifications of standardized geological data products: in terms of spatial structure, each grid cell should have clear geographic coordinate boundaries and a unique identifier; in terms of attributes, the name, unit, detection method and confidence interval of each element should be fully recorded; in terms of file format, it should support GeoTIFF, Grid, Shapefile or NetCDF formats that can be read by mainstream GIS platforms, so as to facilitate subsequent visualization, modeling and sharing.

[0149] More importantly, this dataset inherits the geochemical information from the original high-resolution geological data while adapting to the spatial granularity of low-resolution applications. Whether used for regional anomaly screening, total resource estimation, or as training labels for machine learning models, this synthetic data can be directly put into use, greatly improving the efficiency of geological survey results transformation.

[0150] The above implementation establishes an automated channel for converting refined exploration data into macroscopic geological maps. This technical solution not only utilizes the spatial autocorrelation of geological patterns and the heterogeneity of elemental distribution, but also achieves accurate mapping and fidelity compression between multi-scale data by introducing parameterized configuration and intelligent statistical mechanisms. Compared to traditional methods relying on manual interpretation and experience-based adjustments, this solution significantly improves processing efficiency, enhances the consistency and reproducibility of results, and supports rapid modeling of large-scale areas. The resulting low-resolution geological dataset maintains spatial compatibility with existing geological databases while incorporating the latest high-precision survey results, providing high-quality basic data support for major strategic tasks such as mineral resource prediction, ecological environment monitoring, and land spatial planning.

[0151] Currently, in existing geological exploration practices, the construction of 3D geological data has become a core component of resource assessment and development decision-making. This data is primarily generated through the integration of multiple sources, including geophysical exploration (such as seismic exploration), drilling sampling, and numerical simulation. These data are typically presented in the form of voxels within regular or irregular grids, with each voxel carrying spatial location information and physical properties (such as elemental concentration or density), forming a complex 3D geological dataset.

[0152] However, such datasets face significant challenges in practical applications, especially when dealing with multi-resolution data transformation. The inherent massive scale and dimensional complexity of 3D geological body data make direct modeling operations across the entire space prohibitively expensive. For example, when the dataset size reaches the millions of voxels, common nonlinear mapping methods (such as machine learning prediction algorithms) encounter bottlenecks in memory consumption and computational efficiency, often leading to memory overflows or training non-convergence problems. This is particularly prominent in deep geological exploration, where data sampling density is low and information is incomplete, making it difficult for traditional methods to process efficiently.

[0153] Furthermore, geological bodies typically exhibit sequence stratigraphy and zonation characteristics vertically (such as horizontal bedding of sedimentary strata or vertical elemental gradations in hydrothermal deposits). However, current technologies fail to effectively utilize this natural phenomenon during data conversion. High-resolution shallow data cannot guide predictions for low-resolution deep regions, resulting in a lack of cross-depth knowledge transfer mechanisms. Ultimately, this leads to attribute jumps or distortions in the 3D spatial reconstruction of the conversion results, affecting the overall consistency and geological accuracy of the model. These limitations not only restrict the prediction accuracy for unknown deep regions but also reduce the credibility of the results in resource estimation, necessitating innovative methods for improvement.

[0154] Based on this, this application also discloses a multi-resolution geological data conversion method based on ensemble learning for three-dimensional geological body data.

[0155] Reference Figure 10 A multi-resolution geological data conversion method based on ensemble learning, the specific method includes:

[0156] Step S701: Obtain a three-dimensional geological body dataset;

[0157] The 3D geological dataset is not simply a collection of XYZ coordinate points, but a spatial attribute field generated by fusing various methods such as geophysical exploration (e.g., seismic exploration, gravity / magnetic scanning), drilling sampling analysis, and numerical simulation. Essentially, it is a geological attribute distribution model expressed in a regular or irregular grid format within a 3D Euclidean space. Each data unit, or voxel, carries explicit spatial location information (longitude, latitude, elevation, or depth) and one or more target element concentration values ​​or other physical parameters (e.g., density, porosity, electrical conductivity).

[0158] For example, in oil and gas reservoir evaluation, this type of data may be obtained by combining three-dimensional seismic reflection data volume with downhole logging curve interpolation to form a three-dimensional grid with a resolution of tens of meters; in metal deposit modeling, it may integrate dense borehole test results with structural interpretation results to construct a grade model with mineralization zoning characteristics.

[0159] Step S702: Layer the three-dimensional geological body dataset along the vertical direction to generate multiple two-dimensional slice datasets;

[0160] This step is a crucial dimensionality reduction operation that balances geological rationality and computational efficiency. Although modern computers are capable of processing large-scale three-dimensional tensors, performing complex nonlinear mapping modeling (such as ensemble learning prediction) directly in full three-dimensional space still faces challenges such as high memory consumption, long training cycles, and difficulty in convergence.

[0161] Therefore, this application's embodiments cleverly utilize the layered structure characteristic commonly found in sedimentary rock strata: most geological bodies exhibit a clear sequence of strata in the vertical direction, meaning that the properties within the same stratum change relatively gently, while abrupt changes in lithology or jumps in physical properties may occur between layers. Based on this natural law, "vertical slicing" essentially involves horizontally cutting a three-dimensional body according to elevation or depth intervals, forming a series of two-dimensional planar slices parallel to the Earth's surface, with each layer representing the geological state within a specific depth range. This operation is technically called "horizontal slicing" or "constant-Z sectioning," and is commonly found in the visualization modules of seismic interpretation software (such as Petrel and Kingdom).

[0162] In specific embodiments of this application, layers can be divided according to a fixed spacing set by the user (e.g., one layer every 5 meters) or based on automatic identification of geological interfaces (e.g., based on lithological boundaries or velocity abrupt change surfaces). Each layer slice retains the original spatial grid structure of the XY plane, and aggregates the attribute values ​​of its Z interval into a two-dimensional matrix, thereby transforming the originally complex three-dimensional problem into a series of independent and solvable two-dimensional tasks. More importantly, this layering method respects the temporal sequence logic of geological evolution; each layer can be regarded as a sedimentary product of a certain geological period, allowing subsequent attribute prediction to be carried out in a relatively homogeneous environment, significantly improving the model's learning efficiency and prediction stability.

[0163] Step S703: Input each two-dimensional slice dataset as a low-resolution geological dataset or a high-resolution geological dataset, and independently execute the multi-resolution geological data conversion method as in steps S101 to S105 to obtain the corrected two-dimensional target element data.

[0164] The role of "as low-resolution or high-resolution input" is not a fixed assignment, but rather its functional positioning is dynamically determined based on the actual sampling density and information completeness of each slice. For example, in the shallow near-surface layer, due to the ease of conducting dense sampling (such as soil measurement and trenching), the data usually has high spatial resolution and rich elemental diversity, making it suitable as a high-resolution raw dataset for extracting predictive variables X (such as the content of various associated elements). However, in areas that are difficult to cover by deep drilling, sampling points are sparse and the number of elements detected is limited. In this case, the slice is regarded as a low-resolution target dataset, which needs to rely on the multi-element information provided by the upper layers to predict the missing target elements Y (such as key indicators such as Au and Mo) through a stacking ensemble learning model.

[0165] Understandably, this mechanism essentially achieves cross-depth information transfer and knowledge transfer: shallow, high-precision data becomes the "knowledge teacher" for deep, low-density data, guiding it to complete upscaling reconstruction. The transformation process for each slice runs completely independently, supporting distributed parallel computing and greatly improving overall processing efficiency. This on-demand role-switching strategy avoids the information loss or artificial exaggeration caused by the forced uniformity of sampling standards in traditional methods, truly achieving "locally adapted" intelligent modeling.

[0166] Step S704: All generated two-dimensional target element data are spatially reconstructed along the vertical direction to generate a three-dimensional target element data volume with uniform resolution.

[0167] This step is not merely a formal assembly, but a reconstruction process that ensures geological continuity and spatial consistency. Although each slice has completed its own target element prediction and correction after the previous independent conversion, there may still be issues such as elevation intervals, resolution differences, or abrupt attribute changes between them.

[0168] Therefore, "spatial reconstruction" must include three core technical actions: First, vertical interpolation, which involves filling the intermediate layer between adjacent slices using geostatistical methods (such as ordinary kriging and co-kriging) or machine learning regression models to restore the true trend of attribute gradual change with depth; second, resolution alignment, which unifies the XY grid of all slices to the minimum common granularity (such as 10m×10m) to ensure that there are no misalignments or jagged effects inside the final 3D volume; and finally, 3D volume encapsulation, which reorganizes all slices according to the original Z-axis order and outputs them in a standard 3D raster format (such as VTK, NetCDF, GeoTIFF-3D) for import into professional geological modeling platforms (such as Surpac, Micromine, and Leapfrog) for visualization and resource estimation.

[0169] Most importantly, the process fully considers the vertical evolution of geological bodies. For example, in hydrothermal deposits, ore-forming elements often exhibit a zoning characteristic of being more concentrated at the top and less concentrated at the bottom. If they are simply stacked without transition processing, it will lead to abnormal jumps and distortions. However, by introducing an interpolation algorithm with spatial autocorrelation constraints, this kind of gradual enrichment process can be effectively simulated, making the final model closer to the actual geological background.

[0170] The above implementation achieves high-precision modeling and intelligent resolution upscaling of complex geological bodies under multi-scale conditions. The entire process fully leverages the advantages of modern machine learning in nonlinear relationship mining, while incorporating fundamental geological principles regarding sequence structure, vertical zonation, and spatial continuity, forming a dual support mechanism of "data-driven + knowledge-guided." Compared to traditional methods relying on manual interpolation or empirical extrapolation, this technical solution improves the prediction accuracy and reliability of results for deep, unknown areas, making it particularly suitable for major engineering scenarios requiring detailed 3D characterization, such as mineral prospect prediction, groundwater resource assessment, and carbon sequestration site selection. The final output, a unified-resolution 3D target element data volume, not only possesses a complete spatial topology and physical consistency but can also directly serve reserve calculation, risk assessment, and decision support systems, demonstrating the core value and broad prospects of artificial intelligence technology in promoting the digital transformation of geological science.

[0171] This application also discloses a multi-resolution geological data conversion system based on ensemble learning.

[0172] A multi-resolution geological data conversion system based on ensemble learning, specifically comprising:

[0173] The data acquisition module is used to acquire low-resolution geological datasets containing the target elements, as well as high-resolution geological datasets that do not contain the target elements.

[0174] The spatial aggregation and alignment module is used to perform spatial grid aggregation processing on high-resolution geological datasets to generate a predictor variable matrix that is spatially aligned with low-resolution geological datasets.

[0175] The model training and ensemble module is used to train a Stacking ensemble regression model with the prediction variable matrix as input and the target element values ​​in the low-resolution geological dataset as output. The first layer of base learners outputs the base learner prediction matrix, and the second layer of meta-learners generates the fusion prediction model based on the base learner prediction matrix.

[0176] The high-resolution prediction output module is used to input high-resolution geological datasets into the fusion prediction model and output preliminary predicted values ​​of target elements.

[0177] The spatial error correction module is used to perform spatial error correction on the preliminary predicted values ​​of target elements based on the actual values ​​of target elements in a low-resolution geological dataset, and generate corrected high-resolution target element data.

[0178] This application also discloses a multi-resolution geological data conversion system based on ensemble learning for three-dimensional geological body data.

[0179] A multi-resolution geological data conversion system based on ensemble learning, the conversion system comprising:

[0180] The 3D data acquisition module is used to acquire 3D geological body datasets;

[0181] The slicing module is used to layer a 3D geological body dataset along the vertical direction to generate multiple 2D slice datasets.

[0182] The two-dimensional data correction module is used to take each two-dimensional slice dataset as input as a low-resolution geological dataset or a high-resolution geological dataset, and independently execute steps S101-S105 of the above-mentioned multi-resolution geological data conversion method to obtain the corrected two-dimensional target element data.

[0183] The spatial reconstruction module is used to spatially reconstruct all generated two-dimensional target element data along the vertical direction, generating a three-dimensional target element data volume with uniform resolution.

[0184] The multi-resolution geological data conversion system based on ensemble learning according to the embodiments of this application can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiments.

[0185] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0186] This application also discloses a computer-readable storage medium.

[0187] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the multi-resolution geological data conversion methods based on ensemble learning.

[0188] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0189] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A multi-resolution geological data conversion method based on ensemble learning, characterized in that, The conversion method includes: Obtain a low-resolution geological dataset containing the target element, and a high-resolution geological dataset that does not contain the target element; Spatial grid aggregation is performed on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset; Using the predicted variable matrix as input and the target element values ​​in the low-resolution geological dataset as output, a Stacking ensemble regression model is trained; wherein, the first layer of base learners outputs the base learner prediction matrix, and the second layer of meta-learners generates a fusion prediction model based on the base learner prediction matrix. The high-resolution geological dataset is input into the fusion prediction model, which outputs preliminary predicted values ​​of the target elements. Based on the actual values ​​of the target elements in the low-resolution geological dataset, the regional scale error correction coefficient is calculated. The sum of the actual measurements of the target elements in the low-resolution geological data is used as the true reference quantity. At the same time, the preliminary prediction values ​​generated by the high-resolution model are integrated and summed within the same geographical boundary to obtain the total prediction. The ratio between the two constitutes the error correction coefficient. Based on the spatial grid distribution characteristics of the high-resolution geological dataset, the area weighting factor of each spatial unit is determined. The area weighting factor is directly derived from the geometric properties of each unit. The error correction coefficient and the area weighting factor are applied to the preliminary predicted value of the target element to generate a spatial redistribution result. Integrate the spatial redistribution results of all spatial units and output corrected high-resolution target element data.

2. The multi-resolution geological data conversion method based on ensemble learning according to claim 1, characterized in that, The steps of performing spatial grid aggregation on the high-resolution geological dataset to generate a predictor variable matrix spatially aligned with the low-resolution geological dataset include: Based on the spatial grid structure of the low-resolution geological dataset, define the target grid cell partitioning rules; Based on the target grid cell partitioning rules, the data of each sampling point in the high-resolution geological dataset is mapped to the corresponding target grid cell; Aggregate calculations are performed on all sampled point element values ​​within each target grid cell to generate aggregated feature values ​​for each target grid cell; The aggregated feature values ​​of all target grid cells are integrated to form a predictor variable matrix.

3. The multi-resolution geological data conversion method based on ensemble learning according to claim 2, characterized in that, The steps for training a Stacking ensemble regression model, using the predicted variable matrix as input and the target element values ​​from a low-resolution geological dataset as output, include: Construct a first-layer base learner group containing multiple base learners, and train each base learner separately using the prediction variable matrix and the target element values ​​in the low-resolution geological dataset; The prediction variable matrix is ​​input into each base learner to generate the corresponding base learner prediction results. Integrate the prediction results of all base learners to form a base learner prediction result matrix; Using the prediction result matrix of the base learner as input features and the target element value as the output target, a meta-learner is trained to generate a fusion prediction model.

4. The multi-resolution geological data conversion method based on ensemble learning according to claim 3, characterized in that, The steps of inputting the high-resolution geological dataset into the fusion prediction model and outputting preliminary predicted values ​​of the target elements include: Extracting the spatial unit structure of high-resolution geological datasets; The element data of each spatial unit are independently input into the fusion prediction model; Based on the regression calculation of the fusion prediction model, the predicted value of the target element for each spatial unit is generated; The predicted values ​​of target elements from all spatial units are integrated to form a preliminary set of predicted values ​​for target elements.

5. A multi-resolution geological data conversion method based on ensemble learning according to any one of claims 1 to 4, characterized in that, The conversion method further includes: Obtain a high-resolution geological dataset containing the target elements; Based on the preset low-resolution target grid structure, the high-resolution geological dataset is divided into spatial grid cells. The element values ​​within each spatial grid cell are aggregated and calculated to generate the aggregated value for each spatial grid cell. The aggregated values ​​of all spatial grid cells are integrated to generate a synthetic low-resolution geological dataset.

6. A multi-resolution geological data conversion method based on ensemble learning, characterized in that, The conversion method includes: Obtain a 3D geological body dataset; The three-dimensional geological body dataset is layered vertically to generate multiple two-dimensional slice datasets; Each two-dimensional slice dataset is input as a low-resolution geological dataset or a high-resolution geological dataset, and the multi-resolution geological data conversion method as described in claim 1 is executed independently to obtain the corrected two-dimensional target element data. All generated two-dimensional target element data are spatially reconstructed along the vertical direction to generate a three-dimensional target element data volume with uniform resolution.

7. A multi-resolution geological data conversion system based on ensemble learning, characterized in that, The conversion system includes: The data acquisition module is used to acquire low-resolution geological datasets containing the target elements, as well as high-resolution geological datasets that do not contain the target elements. The spatial aggregation and alignment module is used to perform spatial grid aggregation processing on the high-resolution geological dataset to generate a predictive variable matrix that is spatially aligned with the low-resolution geological dataset. The model training and ensemble module is used to train a Stacking ensemble regression model with the predicted variable matrix as input and the target element values ​​in the low-resolution geological dataset as output targets; wherein, the first layer base learner group outputs the base learner prediction matrix, and the second layer meta learner generates a fusion prediction model based on the base learner prediction matrix. The high-resolution prediction output module is used to input the high-resolution geological dataset into the fusion prediction model and output preliminary predicted values ​​of the target elements. The spatial error correction module is used to perform spatial error correction on the preliminary predicted values ​​of the target elements based on the actual values ​​of the target elements in the low-resolution geological dataset, and generate corrected high-resolution target element data. The spatial error correction module is specifically configured as follows: Based on the actual values ​​of the target elements in the low-resolution geological dataset, the regional scale error correction coefficient is calculated. The sum of the actual measurements of the target elements in the low-resolution geological data is used as the true reference quantity. At the same time, the preliminary prediction values ​​generated by the high-resolution model are integrated and summed within the same geographical boundary to obtain the total prediction. The ratio between the two constitutes the error correction coefficient. Based on the spatial grid distribution characteristics of the high-resolution geological dataset, the area weighting factor of each spatial unit is determined. The area weighting factor is directly derived from the geometric properties of each unit. The error correction coefficient and the area weighting factor are applied to the preliminary predicted value of the target element to generate a spatial redistribution result. Integrate the spatial redistribution results of all spatial units and output corrected high-resolution target element data.

8. A multi-resolution geological data conversion system based on ensemble learning, characterized in that, The conversion system includes: The 3D data acquisition module is used to acquire 3D geological body datasets; The slicing module is used to layer the three-dimensional geological body dataset along the vertical direction to generate multiple two-dimensional slice datasets. The two-dimensional data correction module is used to take each two-dimensional slice dataset as input as a low-resolution geological dataset or a high-resolution geological dataset, and independently execute the multi-resolution geological data conversion method as described in claim 1 to obtain the corrected two-dimensional target element data. The spatial reconstruction module is used to spatially reconstruct all generated two-dimensional target element data along the vertical direction, generating a three-dimensional target element data volume with uniform resolution.

9. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-point geostatistics simulation method based on multi-scale exploration geochemical data

    CN116467855A

  • Determining spatial distributions of petrophysical properties in a subsurface formation

    US20250129701A1