A comprehensive information prospecting prediction method based on multi-source data

CN122595138APending Publication Date: 2026-08-18KUNMING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610754384.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]现有技术在实际应用中仍存在明显不足:多源地学数据中包含大量与矿化无关的背景信息,若未对背景场与异常场进行有效分离,容易使模型受到噪声干扰,降低预测准确性,而且,地学异常并不必然来源于深部矿体,若不能对异常信息进行来源判别并区分矿致异常与非矿致异常,则容易将构造扰动、表生作用或其他地质过程造成的异常误判为找矿指示信息,导致靶区圈定偏差较大

Benefits of technology

通过对多源地学数据进行空间网格化、地学数据体构建以及背景场与异常场分离,能够从原始数据中更精准地提取与成矿相关的异常信息,降低背景噪声对找矿预测结果的干扰,提高输入特征的有效性和纯净度。通过对异常信息进行来源判别,筛选出来源于深部矿体的矿致异常,并将其余异常标记为非矿致异常,通过对比两类异常在地学数据体中的特征差异,筛选出对找矿预测具有指示意义的特征变量类型,使模型能够聚焦于矿致异常与非矿致异常之间的关键区分特征,从而提升对复杂地质条件下异常信息的辨识能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595138A_ABST
    Figure CN122595138A_ABST
Patent Text Reader

Abstract

The application discloses a comprehensive information prospecting prediction method based on multi-source data and belongs to the technical field of mineral exploration, and aims to solve the problem of insufficient reliability of results caused by the diversity of anomaly sources in the covered area and direct prediction without screening. The method comprises the following steps: carrying out multi-source geological data gridding and background field separation in a sample area, obtaining anomaly information and constructing a first feature set; discriminating the sources of the anomaly information, dividing the anomaly into ore-induced anomaly and non-ore-induced anomaly, and determining the types of prospecting characteristic variables through comparative analysis; taking the known ore-bearing and non-ore-bearing grid units as positive and negative samples, extracting characteristic values according to the determined characteristic variable types, and training an integrated prediction model; treating the area to be predicted in the same way and extracting characteristic values, inputting the trained model to obtain the metallogenic probability, and delineating the ore potential grid. Through the collaborative mechanism of source discrimination, feature selection and machine learning, the accuracy of the prospecting prediction result of the covered area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mineral exploration technology, specifically to a comprehensive information-based mineral prospecting prediction method based on multi-source data. Background Technology

[0002] As mineral resource exploration expands into deeper, concealed ore bodies, areas with high vegetation cover and thick sedimentary layers have become important prospecting areas. In these regions, the surface is covered by vegetation or loose sediments, making it difficult to directly obtain mineralization information using traditional geological prospecting methods. The introduction of multi-source remote sensing, deep-penetrating geochemistry, and machine learning technologies has provided new technical approaches for prospecting in covered areas; however, existing technologies still have significant shortcomings in practical applications.

[0003] In existing technologies, CN116823896A discloses a method for predicting target areas in high vegetation cover areas. This method acquires hyperspectral and lidar data, extracts spectral features of the vegetation canopy, and uses a deep learning model to predict the target area range. CN105527230A identifies mineral-induced anomalies using differences in the spectral response of vegetation in the near-infrared band. Both of these methods employ a technical route of "remote sensing data acquisition → vegetation feature extraction → model prediction," relying on the indirect response of vegetation to infer deep mineralization information, but they fail to thoroughly identify the source of surface anomaly signals.

[0004] Existing technologies still have significant shortcomings in practical applications: multi-source geoscientific data contain a large amount of background information unrelated to mineralization. If the background field and the anomaly field are not effectively separated, the model is easily affected by noise, reducing the accuracy of prediction. Moreover, geoscientific anomalies do not necessarily originate from deep ore bodies. If the source of the anomaly information cannot be identified and mineral-induced anomalies cannot be distinguished from non-mineral-induced anomalies, anomalies caused by tectonic disturbances, surface processes, or other geological processes are easily misjudged as mineral exploration indicators, resulting in a large deviation in target area delineation.

[0005] Existing technologies fail to systematically compare and utilize the differences in feature variables between mineral-induced anomalies and non-mineral-induced anomalies during feature selection, making it difficult to effectively distinguish the feature boundaries of the two types of anomalies. This, in turn, affects the stability of the mineralization probability output of the predicted area and the accuracy of the spatial distribution determination.

[0006] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0007] The purpose of this invention is to provide a comprehensive information-based mineral exploration prediction method and system based on multi-source data to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A comprehensive information-based mineral exploration prediction method based on multi-source data includes the following steps: S1: Acquire multi-source geoscience data of the sample area, divide the sample area into spatial grids to obtain several sample grid units, map the multi-source geoscience data to each sample grid unit to obtain geoscience data volume, and perform background field and anomaly field separation processing on the geoscience data volume to obtain anomaly information and construct the first feature set; S2: Distinguish the sources of the abnormal information in the first feature set, determine whether each abnormality originates from a deep ore body, and classify the abnormalities into mineral-induced abnormalities and non-mineral-induced abnormalities; based on the comparative analysis of the geological data volumes of mineral-induced abnormalities and non-mineral-induced abnormalities, determine the types of mineral exploration feature variables. S3: Using known mineral-bearing and mineral-free grid cells within the sample area as positive and negative samples, extract feature values ​​from the corresponding geoscientific data volume according to the mineral exploration feature variable type determined in S2 to form a training input feature set; use the training input feature set and positive and negative sample labels to train the ensemble prediction model and obtain the trained prediction model. S4: Obtain multi-source geoscience data of the area to be predicted, and perform gridding and data mapping in the same way as S1 to obtain geoscience data volumes; according to the feature variable types determined in S2, extract feature values ​​from the geoscience data volumes of each grid unit to be predicted to form the input feature set to be predicted; S5: Input the input feature set to be predicted into the trained prediction model, output the mineralization probability prediction value of each grid cell to be predicted, and determine the grid cells with the mineralization probability prediction value greater than the preset threshold as mineral potential grids, and output the spatial distribution information of the target mineralization potential area.

[0009] Furthermore, the methods for acquiring the geoscientific data volumes described in S1 include: The sample area is divided into regular grids according to a preset grid unit size to obtain the sample grid unit; Multi-source geoscientific data are unified to the same coordinate system. For point sampling data, Kriging interpolation or inverse distance weighting is used to interpolate and assign values ​​to each sample grid cell. For areal vector data, attribute values ​​are assigned according to the geological cell into which the center point of the sample grid cell falls. For raster data, bilinear interpolation or cubic convolution is used to resample to the target grid resolution to obtain the geoscientific data volume of each sample grid cell.

[0010] Furthermore, the construction method of the first feature set in S1 includes: The method employs a multifractal singularity index to calculate the singularity index of each sample grid cell at different observation scales, identifying sample grid cells with singularity indices less than a preset threshold as anomalous information; and / or employs the SA fractal filtering method to convert the gridded data to the frequency domain, constructing a double logarithmic relationship graph between energy spectral density and cumulative area, determining a fractal threshold based on the fractal characteristics in the double logarithmic relationship graph, separating the regional background field and local anomalous field based on the fractal threshold, and identifying sample grid cells corresponding to the local anomalous field as anomalous information; and traversing the anomalous information of all sample grid cells to construct a first feature set.

[0011] Furthermore, the source identification methods in S2 include isotope comparison and element vertical migration simulation. The isotope comparison method includes: collecting ore samples from known deep ore bodies within the sample area and determining their isotopic composition as a reference fingerprint; collecting surface medium samples at the sample grid cell locations corresponding to the anomaly information in the first feature set and determining the compositional characteristics of lead or sulfur isotopes therein; comparing the isotopic composition of the surface medium samples with the reference fingerprint, and determining that the anomaly originates from a deep ore body when the two are consistent within a preset error range; The vertical migration simulation method for elements includes: acquiring geological parameters of the sample area, including overburden thickness, porosity, permeability, and water content; establishing a diffusion-convection model for the migration of target elements from deep ore bodies to the surface, and determining the diffusion coefficient and convection velocity based on the geological parameters; performing forward modeling using the diffusion-convection model to obtain the theoretical surface anomaly distribution; comparing the spatial similarity between the anomaly distribution corresponding to the anomaly information extracted in S1 and the theoretical surface anomaly distribution, and determining that the anomaly information originates from a deep ore body when the similarity is greater than a preset threshold.

[0012] Furthermore, the determination of the mineral exploration feature variable type in S2 includes: extracting the values ​​of each candidate variable from the geoscientific data volume of the grid cell where the mineral-induced anomaly is located and the grid cell where the mineral-induced anomaly is not located, respectively, calculating the statistical difference of each candidate variable between the two types of anomalies; and determining the candidate variable whose statistical difference is greater than a preset threshold as the mineral exploration feature variable.

[0013] Furthermore, the ensemble prediction model described in S3 consists of a random forest learner and a gradient boosting tree learner; The random forest learner is used to handle discrete geological variables, and the gradient boosting tree learner is used to handle continuous geochemical variables. The output of the random forest learner and the output of the gradient boosting tree learner are integrated in a weighted manner, and the weight coefficients of each learner are determined by cross-validation.

[0014] Furthermore, the method for obtaining the trained prediction model described in S3 includes: dividing the training input feature set into a training set and a validation set according to a preset ratio; The training set is used to optimize the parameters of the ensemble prediction model, and the validation set is used to evaluate the model's prediction accuracy and adjust the model's hyperparameters. During training, the optimization objective is to make the model output value approximate the positive and negative sample labels, where the positive sample label is 1 and the negative sample label is 0.

[0015] Further, S4 includes: spatially gridding the region to be predicted according to the same grid cell size and division rules as S1; assigning the multi-source geoscientific data to each grid cell to be predicted according to the same data mapping method as S1; and extracting corresponding feature values ​​from the geoscientific data volume of each grid cell to be predicted according to the feature variable type determined in S2, thereby forming the input feature set to be predicted.

[0016] Furthermore, the step S5, which identifies the predicted grid cells with mineralization probability prediction values ​​greater than a preset probability threshold as mineral potential grids, includes: comparing the mineralization probability prediction values ​​output by the trained prediction model for each predicted grid cell with the preset probability threshold; marking the predicted grid cells with mineralization probability prediction values ​​greater than the preset probability threshold as mineral potential grids, and marking the predicted grid cells with mineralization probability prediction values ​​less than or equal to the preset probability threshold as non-mineral potential grids; and recording the row and column positions of each mineral potential grid within the predicted area to generate a spatial distribution map of the mineral potential grids.

[0017] Furthermore, after generating the spatial distribution map of the mineral potential grid, the method further includes: merging spatially adjacent mineral potential grids into a single mineral potential zone; numbering each mineral potential zone and calculating its area; calculating the mean value of the predicted mineralization probability of each mineral potential grid within each mineral potential zone; classifying each mineral potential zone according to the mean value; and marking the boundary range, number, area, and level of each level of mineral potential zone on the spatial distribution map of the mineral potential grid, as outputting the spatial distribution information of the target mineral potential zone.

[0018] Compared with the prior art, the beneficial effects of the present invention are: By spatially gridding multi-source geoscientific data, constructing geoscientific data volumes, and separating background and anomaly fields, mineralization-related anomaly information can be extracted more accurately from the raw data. This reduces the interference of background noise on mineral exploration prediction results and improves the effectiveness and purity of input features. By identifying the source of anomaly information, mineral-induced anomalies originating from deep ore bodies are selected, while other anomalies are marked as non-mineral-induced anomalies. By comparing the feature differences between the two types of anomalies in the geoscientific data volume, the types of feature variables that are indicative of mineral exploration prediction are selected. This allows the model to focus on the key distinguishing features between mineral-induced and non-mineral-induced anomalies, thereby improving the ability to identify anomaly information under complex geological conditions.

[0019] The integrated prediction model is trained based on known ore-bearing and ore-free samples, enabling the model to learn the mapping relationship between feature variables and mineralization probabilities, thus enhancing its ability to express the mineralization patterns of concealed ore bodies. Since the predicted areas also employ a unified gridding and feature construction method, the predicted mineralization probabilities output by the model exhibit good spatial consistency and comparability, ensuring effective generalization of prediction results from the sample areas to the predicted areas. Furthermore, by combining this with a preset threshold to filter the ore potential grid, the spatial distribution range of the target ore potential area can be more accurately delineated, improving exploration deployment efficiency. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall method flow of the present invention.

[0021] Figure 2 This is a schematic diagram of the entire process of the method of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0023] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0024] Example: Please see Figure 1 and Figure 2 The present invention provides a technical solution: A comprehensive information-based mineral exploration prediction method based on multi-source data includes the following steps: S1: Acquire multi-source geoscience data of the sample area, divide the sample area into spatial grids to obtain several sample grid units, map the multi-source geoscience data to each sample grid unit to obtain geoscience data volume, and perform background field and anomaly field separation processing on the geoscience data volume to obtain anomaly information and construct the first feature set.

[0025] The multi-source geoscientific data includes deep-penetrating geochemical data, geophysical data, remote sensing data, and geological interpretation data. Among them, deep-penetrating geochemical data is used to detect weak geochemical anomalies in the formation of deep concealed ore bodies on the surface, and is obtained through geoelectrochemical extraction or soil fine-particle separation methods. Geophysical data, including gravity and magnetic data, is used to reflect the differences in the physical properties of underground rock masses and structures in the covered area, and is extracted through ground or airborne geophysical measurements. Remote sensing data includes multispectral remote sensing image data and hyperspectral remote sensing image data, which are used to extract tectonic and alteration information related to mineralization, and are acquired through spaceborne or airborne sensors. Geological interpretation data includes stratigraphic and lithological distribution data and fault structure distribution data, which are used to provide spatial distribution information of strata, structures and rock masses in the covered area. This data is obtained through digitization and spatial analysis of existing geological maps.

[0026] S11: Determine the grid unit size based on the sample area's extent and the exploration scale. Divide the sample area into regular grids according to the preset grid unit size to obtain the sample grid units. For example, the sample area is a rectangular area 20km wide east-west and 25km long north-south, with an area of ​​approximately 500km², using a 1:50,000 scale. Based on the general principle that grid unit size in mineral prediction is typically 1 / 50 to 1 / 100 of the exploration area's extent, the preset grid unit size is 100m × 100m. Divide the sample area into a regular grid of 200 rows × 250 columns, resulting in 50,000 sample grid units. Each grid unit is assigned a unique row and column index number in the format (i,j), where i=1,2,…,200, j=1,2,…,250.

[0027] S12: The collected multi-source geoscientific data are uniformly converted to the same coordinate system and elevation datum. In this embodiment, all data adopt the CGCS2000 national geodetic coordinate system and the Gauss-Kruger projection to ensure that various types of data can be accurately matched in spatial location.

[0028] For point-based sampling data, Kriging interpolation or inverse distance weighting is used to interpolate and assign values ​​to each sample grid cell. For area vector data, attribute values ​​are assigned based on the geological cells into which the center point of the sample grid cell falls. For raster data, bilinear interpolation or cubic convolution is used to resample to the target grid resolution.

[0029] After traversing all sample grid cells, each sample grid cell obtains a set of multidimensional geoscientific parameters, which constitute the geoscientific data volume of that sample grid cell.

[0030] S13: Employing the multifractal singularity index method, by calculating the singularity index of each sample grid cell at different observation scales, sample grid cells with singularity indices less than a preset threshold are identified as anomalous information; and / or employing the SA fractal filtering method, the gridded data is converted to the frequency domain, a double logarithmic relationship graph of energy spectral density and cumulative area is constructed, a fractal threshold is determined based on the fractal characteristics in the double logarithmic relationship graph, and the regional background field and local anomalous field are separated based on the fractal threshold, and sample grid cells corresponding to the local anomalous field are identified as anomalous information; the anomalous information of all sample grid cells is traversed and constructed as the first feature set.

[0031] In this embodiment, the singularity index is calculated based on multifractal theory, and the core formula is: ; Where C(ε) is the average elemental concentration within the scale ε window, E is the Euclidean dimension (E is taken as 2 in two-dimensional space), and α is the singularity index. Taking the logarithm of both sides of the formula, we get... During the calculation, for each sample grid cell, a series of square windows with increasing side lengths are defined centered on that cell. The window side lengths are 1, 3, 5, 7, and 9 grid cells respectively, and the average concentration of the specified element within each window is calculated to obtain C(ε) at different scales ε. In a log-log coordinate system, with log(ε) as the abscissa and log[C(ε)] as the ordinate, least squares linear fitting is performed on the data points at each scale to obtain the slope k of the fitted line. The singularity index can be obtained. When α < 2, it indicates that the element concentration around the sample grid cell shows a local enrichment trend, and sample grid cells with α less than a preset threshold (such as α < 1.8) are identified as abnormal information.

[0032] When using the SA fractal filtering method to identify anomalous information, the gridded data is converted to the frequency domain, and the relationship between the energy spectral density and the cumulative area is analyzed in a double logarithmic coordinate system. The fractal threshold is determined by identifying the boundary points of different slope segments. A filter is constructed to separate the frequency domain data into the regional background field and the local anomalous field. The local anomalous field components are converted back to the spatial domain through inverse Fourier transform, and the sample grid cells corresponding to the local anomalous field are identified as anomalous information.

[0033] The SA fractal filtering method separates the background field from the anomaly field in the frequency domain based on the fractal relationship between energy spectral density and cumulative area. First, the geophysical or geochemical data of each sample grid cell are converted from the spatial domain to the frequency domain using a two-dimensional Fourier transform, and the energy spectral density S at each frequency point is calculated. A logA-logS relationship graph is plotted in a double logarithmic coordinate system with energy spectral density S as the abscissa and the cumulative area A of frequencies with energy spectral density greater than that as the ordinate. The graph typically presents as straight line segments with different slopes, representing the fractal characteristics corresponding to different geological processes. A fractal threshold S0 is determined by identifying the boundary points between the straight line segments. A low-pass filter (S < S0) and a high-pass filter (S ≥ S0) are constructed using S0 as the boundary to separate the frequency domain data into two components: a regional background field and a local anomaly field. The low-frequency component after low-pass filtering corresponds to the regional background field, and the high-frequency component after high-pass filtering corresponds to the local anomaly field. The local anomalous field components are transformed back to the spatial domain through inverse Fourier transform to obtain a local anomalous spatial distribution map, and the sample grid cells corresponding to the local anomalous field are identified as anomalous information.

[0034] Traverse all sample grid cells and summarize the abnormal information identified by any one or a combination of the two methods mentioned above. Each sample grid cell identified as abnormal information corresponds to an abnormal record, and the set of all abnormal records constitutes the first feature set.

[0035] S2: Distinguish the sources of the abnormal information in the first feature set, determine whether each abnormality originates from a deep ore body, and classify the abnormalities into mineral-induced abnormalities and non-mineral-induced abnormalities; based on the comparative analysis of the geological data volumes of mineral-induced abnormalities and non-mineral-induced abnormalities, determine the types of mineral exploration feature variables.

[0036] The methods for determining the source include isotope comparison and elemental vertical migration simulation. The isotope comparison method is based on the principle that deep ore bodies and overburden containment materials have different isotopic "fingerprints" for source tracing. Deep, concealed ore bodies develop specific isotopic composition characteristics during mineralization, while exogenous materials in the overburden (such as aeolian sand and bioclastic debris) do not possess these characteristics. Therefore, if the isotopic composition of anomaly samples from the surface is consistent with that of deep ore bodies, it can be determined that the anomalous material originates from a deep ore body.

[0037] In this embodiment, firstly, ore samples are collected from known deep ore bodies within the sample area, and their isotopic composition is determined as a baseline fingerprint. The determination indicators are typically... , and Three ratios, or the sulfur isotope composition determined by stable isotope mass spectrometry, with the determination index being... Value, i.e., in the sample and The ratio of the above measurements to the international standard VCDT is used as the isotopic "baseline fingerprint" of the deep ore body in the mining area, and a baseline database is established.

[0038] Subsequently, surface medium samples were collected at the sample grid cell locations corresponding to the abnormal information in the first feature set. The collected surface medium samples were analyzed using the same analytical methods and instrument conditions as the ore samples to determine the compositional characteristics of lead isotopes or sulfur isotopes.

[0039] The isotopic composition of a surface medium sample is compared with the reference fingerprint. When the two are consistent within a preset error range, the anomaly is determined to originate from a deep ore body. For example, taking lead isotopes as an example, the isotopic composition of the surface sample and the ore sample is calculated... , and The relative differences among the three ratios are considered. If all three ratios are within the preset error range, the surface sample and the deep ore body are deemed to have the same origin, meaning the anomaly originates from a deep, concealed ore body. If any ratio difference exceeds the error range, it indicates that the surface anomaly is not caused by a deep ore body, but rather by interference from the overburden. It should be noted that the preset error range is typically ±0.3% to ±0.5%, depending on the instrument's measurement accuracy.

[0040] It should be noted that the applicability of this method varies depending on the mining area. Before implementing S2, surface samples from the known ore body directly above it and from the background area away from it can be collected to assess the applicability of the isotopic composition. If there is a statistically significant difference in the isotopic ratio between the samples above the ore body and the background area, it indicates that isotopic tracing is effective in this mining area; if there is no significant difference, it indicates that the isotopic comparison method is not suitable for this mining area, and the vertical migration simulation method should be used instead.

[0041] The vertical migration simulation method of elements is based on the principle that ore-forming elements in deep ore bodies can migrate upwards along channels such as micro-fractures and pores in the overburden layer and form observable anomalous distributions near the surface to determine their origin.

[0042] In this embodiment, the geological parameters of the sample area are first obtained, including the thickness of the overburden, porosity, permeability and water content; wherein the thickness of the overburden is obtained by drilling or geophysical exploration, the porosity is obtained by borehole core laboratory measurement or geophysical logging, the permeability is obtained by borehole pumping test or core permeability test, and the water content is obtained by indoor geotechnical test of borehole core sample.

[0043] Based on the above parameters, a diffusion-convection model for the migration of target elements from deep ore bodies to the surface is established. This model is a one-dimensional vertical partial differential equation, and its expression is as follows: Where C is the target element concentration in mg / kg or ppm, t is time in years, z is the depth coordinate with the surface as the origin and downward as the positive direction in meters, and D is the diffusion coefficient in meters. 2 / s, v is the convection velocity in m / s, and λ is the first-order attenuation constant in s. -1 The left side of the equation The first term on the right-hand side of the equation represents the rate of change of element concentration over time. This represents the concentration change caused by molecular diffusion, describing the migration of elements from high-concentration regions to low-concentration regions driven by a concentration gradient. The second term... This indicates the concentration change caused by groundwater convection, describing the migration process of elements as they move with the overall flow of groundwater. (The third item...) This represents the attenuation loss of elements during migration. The equation comprehensively reflects the control of vertical element migration by diffusion, convection, and attenuation mechanisms, and describes the variation of element concentration with time and depth.

[0044] Key parameters in the model include the diffusion coefficient D and the convection velocity v. The diffusion coefficient D describes the molecular diffusion ability of the target element in the porous medium of the capping layer. The diffusion coefficient D can be determined based on the porosity and water content, combined with the molecular diffusion coefficient of the target element in water. Its calculation formula is as follows: in, φ is the molecular diffusion coefficient of the target element in free water, which is a constant and obtained by looking up a table; φ is the porosity and θ is the water content.

[0045] The convection velocity *v* describes the average seepage velocity of groundwater in the pores of the overburden layer. It is quantitatively calculated based on Darcy's law. First, the Darcy velocity *μ* is obtained by multiplying the permeability *k* by the hydraulic gradient *J*. Then, the Darcy velocity *μ* is divided by the porosity *φ* to correct for the true average convection velocity of the fluid within the pores, which is the convection velocity *v* of elements migrating with groundwater in this embodiment. The calculation formula is as follows: The permeability k reflects the water permeability of the rock mass, and the hydraulic gradient refers to the head loss per unit distance along the seepage path, which is usually obtained by the following method: Among the existing long-term hydrogeological observation wells in the sample area, two wells located in the main seepage direction are selected, and their stable water level elevations are measured to obtain the head difference ΔH. The seepage path distance L between the two wells is then measured. Calculated.

[0046] The decay constant λ describes the rate at which the concentration of a target element decreases during migration due to adsorption, precipitation, radioactive decay, and other processes. Its value can be determined through laboratory batch adsorption tests or dynamic soil column tests, or by referring to empirical values ​​in existing literature on similar covered areas.

[0047] Boundary conditions are used to define the physical state of the governing equations at the upper and lower boundaries of the computational domain, ensuring that the equations have a unique solution. In this embodiment, the lower boundary is set at the top of the deep ore body, and the boundary type is a first-type boundary condition (Dirichlet boundary condition). The value is the known target element concentration of the ore body, indicating that the ore body continuously provides ore-forming elements to the overburden as a stable material source. The upper boundary is set at the surface, and the boundary type is a third-type boundary condition (Robin boundary condition), expressed as surface concentration flux, reflecting the exchange process of elements between the surface and the atmosphere or surface water.

[0048] The theoretical surface anomaly distribution is obtained by forward modeling using the aforementioned diffusion-convection model. This embodiment employs the finite difference method: the computational domain is the overburden thickness, and the spatial domain is vertically discretized into several equally spaced nodes. The time domain is calculated progressively from the mineralization period to the present. Within each time step, the partial differential terms in the governing equations are approximated using a difference scheme, and the concentration values ​​for the next time step are recursively calculated from the concentration values ​​of each node at the previous time step. After calculation across the entire time domain, the concentration values ​​of the surface nodes at the current time step are extracted, representing the theoretical surface anomaly distribution, which serves as the benchmark for subsequent similarity comparisons.

[0049] Subsequently, the spatial similarity comparison between the anomaly distribution corresponding to the anomaly information extracted in S1 and the theoretical surface anomaly distribution is performed. In this embodiment, taking the anomaly grid cell extracted in S1 as the center, the normalized elemental concentration value within its preset spatial neighborhood is extracted as the actual anomaly distribution vector. At the same time, the normalized concentration value at the same spatial location in the theoretical anomaly distribution is extracted as the theoretical anomaly distribution vector. It should be noted that the normalized elemental concentration value is a standardized value obtained by mapping the original geochemical elemental concentration data of each grid cell in the target exploration area to the interval of 0 to 1 using the extreme value normalization method. After normalization, the differences in elemental concentration magnitude, background baseline interference, and dimensional influence can be eliminated, making the concentration data of different spatial locations and different elements horizontally comparable, which is convenient for constructing anomaly distribution vectors and carrying out quantitative calculation of spatial similarity.

[0050] The spatial correlation coefficient between the actual anomaly distribution vector and the theoretical anomaly distribution vector is calculated as a quantitative indicator of similarity.

[0051] When the spatial correlation coefficient is greater than the preset threshold, it indicates that the actual anomaly distribution is highly consistent with the theoretical anomaly distribution in terms of spatial morphology, and the anomaly information is determined to originate from a deep concealed ore body; when the correlation coefficient is less than or equal to the preset threshold, it indicates that the actual anomaly distribution cannot be explained by the migration of deep ore bodies, and is determined to be caused by non-mineralized processes such as overburden interference.

[0052] After identifying the source of all anomaly information in the first feature set using the aforementioned isotope comparison method or elemental vertical migration simulation method, the sample grid cells corresponding to the anomalies determined to originate from deep ore bodies are denoted as mineral-induced anomaly grid cells, and the remaining sample grid cells are denoted as non-mineral-induced anomaly grid cells. Based on the comparative analysis of the geoscientific data volumes of the mineral-induced anomaly grid cells and the non-mineral-induced anomaly grid cells, the values ​​of each candidate variable in the geoscientific data volume of the grid cells containing the two types of anomalies are extracted respectively. The statistical difference of each candidate variable between the two types of anomalies is calculated, and the candidate variables with statistical differences greater than a preset threshold are determined as mineral exploration feature variable types for use in S3 model training.

[0053] For example, within a sample area, there are 100 anomalous grid cells, of which 20 are labeled as mineral-induced anomalies and 80 as non-mineral-induced anomalies. Candidate variables are extracted from the geoscientific data volumes of all anomalous grid cells, including: the singularity index of Cu, the concentration of Pb, the distance from the fault zone, the topographic slope, and the remotely sensed alteration intensity. The average values ​​of each candidate variable are calculated for the mineral-induced anomaly group and the non-mineral-induced anomaly group. It is found that the differences in the singularity index of Cu, the distance from the fault zone, and the remotely sensed alteration intensity between the mineral-induced and non-mineral-induced anomaly groups are greater than preset thresholds, while the differences in topographic slope and Pb concentration are less than preset thresholds. Therefore, the singularity index of Cu, the distance from the fault zone, and the remotely sensed alteration intensity are identified as mineral exploration feature variables, while topographic slope and Pb concentration are not included in the mineral exploration feature variable types.

[0054] S3: Using known ore-bearing grid units within the sample area as positive samples and known non-ore-bearing grid units as negative samples, for each positive and negative sample grid unit, feature values ​​are extracted one by one from the geoscientific data volume of the corresponding grid unit according to the mineral exploration feature variable type determined in S2. Each grid unit forms a feature vector. The set of feature vectors of all positive and negative samples constitutes the training input feature set. The training input feature set is used as the model input, and the positive and negative sample labels are used as the model fitting target. The integrated prediction model composed of multiple different learners is trained to obtain the trained prediction model.

[0055] During training, the feature values ​​of each sample in the training set are used as model input, and the positive and negative labels of the samples are used as the model fitting targets, where positive labels are assigned a value of 1 and negative labels are assigned a value of 0. The model outputs a predicted value between 0 and 1. The error between the predicted value and the label value is calculated using a loss function. Optimization algorithms such as backpropagation or gradient descent are used to iteratively adjust the model's internal parameters, gradually bringing the predicted value closer to the label value. The training process continues until the model's error on the training set decreases below a preset standard and its prediction accuracy on the validation set stabilizes.

[0056] The ensemble prediction model consists of two learners: a random forest learner and a gradient boosting tree learner. Their outputs are combined through a weighted ensemble. The purpose of choosing two different learners is twofold: the random forest learner, by ensembled multiple decision trees in parallel, can effectively handle discrete geological variables such as stratigraphic lithology encoding and structural density; the gradient boosting tree learner, through sequential iteration to gradually approximate the true labels, excels at capturing high-order nonlinear relationships between continuous geochemical variables such as the Cu singularity index and local magnetic anomalies after SA filtering, and mineralization probabilities. The two learners complement each other, improving the model's adaptability and prediction accuracy to multi-source heterogeneous geoscientific data in this case.

[0057] It should be noted that the random forest learner consists of multiple decision trees built in parallel, each independently constructed for different subsets of samples drawn from the training set. In a specific embodiment, the number of decision trees is preset to 200, the maximum depth of a single tree is preset to 15 layers, and the minimum number of samples per leaf node is set to 5. Each decision tree starts from the root node and randomly selects some features as candidates from the mineral exploration feature variables determined in S2. For example, at a certain node, three features may be randomly selected as candidates: Cu element singularity index, distance from the contact zone of the rock mass, and remote sensing hydroxyl alteration intensity. The optimal feature and splitting threshold are selected according to the criterion of minimum Gini impurity for node splitting. Through this randomization at the feature level, different decision trees focus on different combinations of mineral-controlling factors during the construction process. For example, some trees focus on the relationship between geochemical anomalies and tectonic distance, while others focus on the correlation between remote sensing alteration and the contact zone of the rock mass, thereby comprehensively evaluating the mineralization potential of the grid unit from multiple perspectives. During prediction, the arithmetic mean of the mineralization probability output values ​​of all decision trees for the same grid cell is taken, and this average value is used as the mineralization probability prediction value of the random forest learner.

[0058] It should be noted that the gradient boosting tree learner adopts a boosting serial iterative architecture. Its training process is as follows: First, the proportion of positive samples in the training samples is used as the initial prediction value, reflecting the overall mineral content benchmark of the sample area. Then, multiple iterations are performed. In each iteration, the residual between the current model's predicted mineralization probability and the true label (1 or 0) is calculated, and a new decision tree is created specifically to fit this residual. A positive residual indicates that the current model underestimates the mineralization probability of that grid cell, while a negative residual indicates that the current model overestimates or underestimates the mineralization probability of that grid cell. The new tree corrects the model's prediction bias by fitting these residuals. For example, if the true label of a certain mineral-causing anomaly grid cell is 1, and the current model's predicted probability is only 0.6, the residual is +0.4. The new tree will learn to address this bias, thereby improving the predicted probability of that grid cell. After the new tree is fitted, its output is multiplied by the learning rate and added to the model output. L1 and L2 regularization terms are introduced to constrain the weights of the leaf nodes of each tree, preventing the model from overfitting to the local noise features of individual grid cells in the training set, thus ensuring stable generalization ability under complex geological conditions. After training, the initial baseline output is accumulated with the weighted outputs of the newly built decision trees in all iterations to obtain the mineralization probability prediction value of the gradient boosting tree learner for each grid cell.

[0059] The training input feature set is divided into a training set and a validation set according to a preset ratio; in this embodiment, the ratio is 7:3. The training set is used to input each base learner for internal model parameter fitting and weight updates; the validation set uses independent samples that were not involved in training to test the model's generalization ability. The validation set uses the prediction accuracy and AUC value as evaluation indicators to guide the iterative optimization of the model's hyperparameters.

[0060] For the random forest learner, with the goal of optimizing the validation set evaluation index, we iterate through the preset parameter value combinations: successively adjust the number of decision trees and the maximum depth of a single decision tree, while keeping the other parameters unchanged. Under each set of parameters, we complete the model training and validation set accuracy evaluation, and retain the parameter combination with the best validation effect as the final hyperparameters.

[0061] For gradient boosting tree learners, the same grid parameter optimization method is used, traversing the combination of learning rate and maximum depth of a single tree within a preset value range; after each parameter update, the model is retrained and the validation set evaluation index is calculated, and the parameter configuration that makes the validation set prediction accuracy the highest and does not cause overfitting is selected, thus completing the hyperparameter solidification.

[0062] When the evaluation metric of the validation set no longer improves after multiple iterations, and the difference in accuracy between the training set and the validation set is less than a set threshold, the model parameter optimization is considered complete, hyperparameter optimization is stopped, and the final structure and parameters of the ensemble prediction model are determined.

[0063] The outputs of the trained random forest learner and gradient boosting tree learner are integrated using a weighted method to obtain the final mineralization probability prediction. The weight coefficients of each learner are determined through cross-validation: the training set is further divided into multiple subsets, and the performance of the two learners is trained and evaluated on different subsets. The weight combination that minimizes the overall prediction error is selected as the final ensemble weights. In this embodiment, five-fold cross-validation is used, with a weight search range of 0.1 to 0.9 and a step size of 0.1.

[0064] After the above training and weight determination, the trained prediction model is obtained and used in step S5 to predict the mineralization probability of the area to be predicted.

[0065] S4: Obtain multi-source geoscientific data for the region to be predicted. The data types are consistent with those of the sample region in S1, including deep-penetrating geochemical data, geophysical data, remote sensing data, and geological interpretation data. Gridding and data mapping are performed in the same manner as in S1 to obtain geoscientific data volumes. Feature values ​​are extracted from the geoscientific data volumes of each grid cell to be predicted, according to the feature variable types determined in S2, to form the input feature set for prediction.

[0066] Based on the mineral exploration characteristic variable types determined in S2, corresponding feature values ​​are extracted one by one from the geoscientific data volume of each grid cell to be predicted, forming the input feature set to be predicted. The list of characteristic variable types was determined in S2 through comparative analysis of mineral-induced anomalies and non-mineral-induced anomalies in the sample area. The S4 stage strictly follows this list, without adding, deleting, or changing any characteristic variable types.

[0067] It should be noted that all processing parameters in step S4 that are consistent with those in S1, and the list of feature variable types that are consistent with those in S2, were determined during the sample region modeling phase of S1-S3. This consistent design ensures that after the data in the region to be predicted undergoes feature engineering pipeline processing identical to that of the training data, the resulting input feature set is highly consistent with the training input feature set in terms of feature dimension, numerical units, and statistical distribution. This enables the trained prediction model to effectively predict the mineralization probability of the region to be predicted.

[0068] S5: Input the input feature set to be predicted into the trained prediction model. This model has learned the mapping law from multi-source geological feature variables to mineralization probability in the sample area. After receiving the input feature set to be predicted, it calculates the mineralization probability prediction value for each grid cell to be predicted, and outputs the mineralization probability prediction value corresponding to each grid cell. The value is between 0 and 1. The closer it is to 1, the higher the probability that there is a hidden ore body in the grid cell.

[0069] Grid cells with predicted mineralization probabilities greater than a preset probability threshold are identified as mineral potential grids. The preset probability threshold is determined in step S3 using the optimal prediction accuracy value on the validation set; in this embodiment, it is set to 0.7. Grid cells with predicted mineralization probabilities greater than the preset probability threshold are marked as mineral potential grids, while those with predicted mineralization probabilities less than or equal to the preset probability threshold are marked as non-mineral potential grids. The row, column, and line positions of each mineral potential grid within the prediction area are recorded, generating a spatial distribution map of the mineral potential grids. Mineral potential grids and non-mineral potential grids are identified in the map using specific colors or symbols.

[0070] After generating the spatial distribution map of mineral potential grids, the discrete mineral potential grids are further integrated into contiguous mineralized potential areas for field exploration deployment. The eight-neighbor connectivity method is used to merge spatially adjacent mineral potential grids into a single mineralized potential area. Each mineralized potential area is sequentially numbered, and its area is calculated by multiplying the number of mineral potential grids within the area by the area of ​​a single grid cell. The arithmetic mean of the predicted mineralization probability values ​​of all mineral potential grids within each mineralized potential area is calculated. All mineralization potential areas are sorted in descending order of their mean mineralization probability values ​​and divided into high, medium, and low potential levels based on their ternary thirds. The first ternary third is designated as Level 1 target areas, the middle ternary third as Level 2 target areas, and the last ternary third as Level 3 target areas. The boundary range, number, area, and level of each level of mineralized potential area are marked on the mineral potential grid spatial distribution map, serving as the output spatial distribution information of the target mineralized potential areas.

[0071] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0072] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0073] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0074] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A comprehensive information-based mineral exploration prediction method based on multi-source data, characterized in that, Includes the following steps: S1: Acquire multi-source geoscience data of the sample area, divide the sample area into spatial grids to obtain several sample grid units, map the multi-source geoscience data to each sample grid unit to obtain geoscience data volume, and perform background field and anomaly field separation processing on the geoscience data volume to obtain anomaly information and construct the first feature set; S2: Distinguish the sources of the abnormal information in the first feature set, determine whether each abnormality originates from a deep ore body, and classify the abnormalities into mineral-induced abnormalities and non-mineral-induced abnormalities; based on the comparative analysis of the geological data volumes of mineral-induced abnormalities and non-mineral-induced abnormalities, determine the types of mineral exploration feature variables. S3: Using known mineral-bearing and mineral-free grid cells within the sample area as positive and negative samples, extract feature values ​​from the corresponding geoscientific data volume according to the mineral exploration feature variable type determined in S2 to form a training input feature set; The ensemble prediction model is trained using the training input feature set and positive and negative sample labels to obtain the trained prediction model. S4: Obtain multi-source geoscience data of the area to be predicted, and perform gridding and data mapping in the same way as S1 to obtain geoscience data volumes; according to the feature variable types determined in S2, extract feature values ​​from the geoscience data volumes of each grid unit to be predicted to form the input feature set to be predicted; S5: Input the input feature set to be predicted into the trained prediction model, output the mineralization probability prediction value of each grid cell to be predicted, and determine the grid cells with the mineralization probability prediction value greater than the preset threshold as mineral potential grids, and output the spatial distribution information of the target mineralization potential area.

2. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The methods for acquiring the geoscientific data volumes described in S1 include: The sample area is divided into regular grids according to a preset grid unit size to obtain the sample grid unit; Multi-source geoscientific data are unified to the same coordinate system. For point sampling data, Kriging interpolation or inverse distance weighting is used to interpolate and assign values ​​to each sample grid cell. For areal vector data, attribute values ​​are assigned according to the geological cell into which the center point of the sample grid cell falls. For raster data, bilinear interpolation or cubic convolution is used to resample to the target grid resolution to obtain the geoscientific data volume of each sample grid cell.

3. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The construction methods of the first feature set in S1 include: The method employs a multifractal singularity index to calculate the singularity index of each sample grid cell at different observation scales, identifying sample grid cells with singularity indices less than a preset threshold as anomalous information; and / or employs the SA fractal filtering method to convert the gridded data to the frequency domain, constructing a double logarithmic relationship graph between energy spectral density and cumulative area, determining a fractal threshold based on the fractal characteristics in the double logarithmic relationship graph, separating the regional background field and local anomalous field based on the fractal threshold, and identifying sample grid cells corresponding to the local anomalous field as anomalous information; and traversing the anomalous information of all sample grid cells to construct a first feature set.

4. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The methods for source identification in S2 include isotope comparison and element vertical migration simulation. The isotope comparison method includes: collecting ore samples from known deep ore bodies within the sample area and determining their isotopic composition as a reference fingerprint; collecting surface medium samples at the sample grid cell locations corresponding to the anomaly information in the first feature set and determining the compositional characteristics of lead or sulfur isotopes therein; comparing the isotopic composition of the surface medium samples with the reference fingerprint, and determining that the anomaly originates from a deep ore body when the two are consistent within a preset error range; The vertical migration simulation method for elements includes: acquiring geological parameters of the sample area, including overburden thickness, porosity, permeability, and water content; establishing a diffusion-convection model for the migration of target elements from deep ore bodies to the surface, and determining the diffusion coefficient and convection velocity based on the geological parameters; performing forward modeling using the diffusion-convection model to obtain the theoretical surface anomaly distribution; comparing the spatial similarity between the anomaly distribution corresponding to the anomaly information extracted in S1 and the theoretical surface anomaly distribution, and determining that the anomaly information originates from a deep ore body when the similarity is greater than a preset threshold.

5. The mineral exploration prediction method based on multi-source data according to claim 4, characterized in that: The determination of the mineral exploration feature variable type in S2 includes: extracting the values ​​of each candidate variable from the geoscientific data volume of the grid cell where the mineral-induced anomaly is located and the grid cell where the mineral-induced anomaly is not located, respectively, calculating the statistical difference of each candidate variable between the two types of anomalies; and determining the candidate variable whose statistical difference is greater than a preset threshold as the mineral exploration feature variable.

6. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The ensemble prediction model described in S3 consists of a random forest learner and a gradient boosting tree learner. The random forest learner is used to handle discrete geological variables, and the gradient boosting tree learner is used to handle continuous geochemical variables. The output of the random forest learner and the output of the gradient boosting tree learner are integrated in a weighted manner, and the weight coefficients of each learner are determined by cross-validation.

7. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The method for obtaining the trained prediction model described in S3 includes: dividing the training input feature set into a training set and a validation set according to a preset ratio; The training set is used to optimize the parameters of the ensemble prediction model, and the validation set is used to evaluate the model's prediction accuracy and adjust the model's hyperparameters. During training, the optimization objective is to make the model output value approximate the positive and negative sample labels, where the positive sample label is 1 and the negative sample label is 0.

8. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: S4 includes: The region to be predicted is spatially gridded according to the same grid cell size and division rules as in S1; the multi-source geoscientific data is assigned to each grid cell to be predicted according to the same data mapping method as in S1; according to the feature variable type determined in S2, the corresponding feature values ​​are extracted from the geoscientific data volume of each grid cell to be predicted to form the input feature set to be predicted.

9. The mineral exploration prediction method based on multi-source data according to claim 1, characterized in that: The method described in S5 for determining the predicted grid cells with a mineralization probability prediction value greater than a preset probability threshold as mineral potential grid cells includes: comparing the mineralization probability prediction value output by the trained prediction model for each predicted grid cell with the preset probability threshold one by one. The predicted grid cells with a mineralization probability prediction value greater than the preset probability threshold are marked as mineral potential grids, and the predicted grid cells with a mineralization probability prediction value less than or equal to the preset probability threshold are marked as non-mineral potential grids. Record the row and column positions of each mineral potential grid within the area to be predicted, and generate a spatial distribution map of the mineral potential grid.

10. A comprehensive information-based mineral exploration prediction method based on multi-source data according to claim 9, characterized in that: After generating the spatial distribution map of mineral potential grids, the method further includes: merging spatially adjacent mineral potential grids into a single mineral potential zone; numbering each mineral potential zone and calculating its area; calculating the mean value of the predicted mineralization probability of each mineral potential grid within each mineral potential zone; classifying each mineral potential zone according to the mean value; and marking the boundary range, number, area, and level of each level of mineral potential zone on the spatial distribution map of the mineral potential grids, as outputting the spatial distribution information of the target mineral potential zone.

Citation Information

Patent Citations

  • Remote sensing prospecting method in high-vegetation-coverage area based on near infrared wave bands in range of 670 to 680 nm

    CN105527230A