A method for constructing a mineral resource prediction agent based on a large model

By constructing a large-scale model of mineral resource prediction intelligent agent, and utilizing causal knowledge graphs and unified feature vectors for mineralization processes, the problems of difficulty in characterizing the causal structure of mineralization processes and data gaps in mineral resource prediction are solved. This optimizes the active sampling task and improves the efficiency of exploration resource allocation.

CN121543696BActive Publication Date: 2026-03-20JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610069929.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-03-20
Estimated Expiration
2046-01-20

AI Technical Summary

Technical Problem

Existing mineral resource prediction methods struggle to explicitly characterize the causal structure of mineralization processes and lack support from proactive sampling driven by data gaps, thus affecting the efficiency of exploration resource allocation.

Method used

We construct a mineral resource prediction agent based on a large model. By using a causal knowledge graph of the mineralization process and a unified feature vector, we infer the mineralization process chain, calculate data gap indicators, and generate active sampling tasks. We then optimize the sampling scheme by combining sampling cost and information gain.

Benefits of technology

It enables explicit representation and quantitative characterization of mineralization event chains and key control nodes, improves the overall efficiency of sampling schemes in terms of cost and benefits, and optimizes the quantitative trade-off between exploration resource input and information gain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543696B_ABST
    Figure CN121543696B_ABST
Patent Text Reader

Abstract

The application discloses a kind of mineral resources prediction intelligent agent construction methods based on large model, it is related to large model application technical field, including, obtain the geological data and album of target mineral research area, utilize large model mineralization process text structured extraction to construct mineralization process causal knowledge diagram, and geoscience data is rasterized and feature coding is the uniform feature vector of mineralization node;Under the control of large model, the link of mineralization process is reasoned, the support and importance of mineralization node are calculated, and data gap index and active sampling task are generated accordingly;Expected information gain is evaluated by mixed density network and bayesian update, and the scheme is optimized under the constraint of sampling cost and iteratively updated representation.The present application realizes quantitative leak detection and filling for key information missing of mineralization;Also realize the quantitative description of mineralization process and the balance of exploration resource input and information gain.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large model application, and in particular to a large model-based mineral resource prediction agent construction method. BACKGROUND

[0002] Mineral resource prediction is a key link in regional metallogenic regularity research and ore prospecting deployment. Traditional work usually relies on geological, geophysical and geochemical exploration and remote sensing multi-source geoscience data, combined with metallogenic model and expert experience to construct ore prediction model. Common methods include evidence weight, logistic regression, random forest and convolutional neural network, which quantize geological bodies, structures, physical property anomalies and geochemical anomalies into spatial features, evaluate the metallogenic favorability of the study area and divide the ore prospecting target area. It plays an important role in improving the utilization efficiency of geoscience data and supporting mineral resource exploration decision-making.

[0003] However, in terms of mechanism expression and data acquisition collaboration, the conventional mineral resource prediction method still has deficiencies. On the one hand, most models focus on using spatial correlation features for metallogenic evaluation, and it is difficult to explicitly represent the causal chain and key control nodes of the metallogenic event. On the other hand, the prediction results lack quantitative linkage with subsequent geological mapping, geophysical exploration and drilling, and it is difficult to identify data gaps based on missing key metallogenic information and plan active sampling tasks, affecting the efficiency of resource allocation. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a large model-based mineral resource prediction agent construction method to solve the problem that the existing technology cannot explicitly depict the causal structure of the metallogenic process and lacks support for active sampling driven by data gaps.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] The present application provides a large model-based mineral resource prediction agent construction method, which comprises: obtaining geological data and atlas of the target mineral research area, using a large model to structureally extract the metallogenic process text, constructing a metallogenic process causal knowledge graph, and rasterizing and feature encoding multi-source geoscience data to establish a unified feature vector corresponding to the metallogenic nodes in the metallogenic process causal knowledge graph;

[0008] Under the control of the large model, based on the metallogenic process causal knowledge graph and the unified feature vector, the metallogenic process link is inferred, and based on the metallogenic process link, the support and importance of the metallogenic nodes in the metallogenic process link are calculated;

[0009] According to the support and importance of the ore-forming node, a data gap index is generated, the sampling type and spatial position required to make up the data gap are derived by the large model, an active sampling task is generated, and optimization is performed in combination with the sampling cost and expected information gain to obtain an active sampling scheme;

[0010] New geoscience data generated by active sampling is received, the unified feature vector is updated, the data gap index is recalculated under the control of the large model and is iteratively updated.

[0011] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the ore-forming process causal knowledge graph comprises,

[0012] Collect data sources, and obtain geological data and atlases of the target mineral research area from the data sources;

[0013] The geological data and atlases of the target mineral research area are cleaned, de-duplicated, and format-normalized;

[0014] The large model is adaptively trained based on the processed geological data and atlases of the target mineral research area;

[0015] The ore-forming process text is input into the adaptively trained large model, structured extraction is performed through a prompt template, and ore-forming nodes, ore-forming causal relationships, and ore-forming time sequence relationships are obtained;

[0016] The ore-forming nodes, ore-forming causal relationships, and ore-forming time sequence relationships are stored in a directed graph structure, and an ore-forming process causal knowledge graph is constructed;

[0017] The ore-forming nodes in the ore-forming process causal knowledge graph are used to represent ore-forming events, and the directed edges are used to represent ore-forming causal relationships.

[0018] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the establishment of a unified feature vector corresponding to the ore-forming nodes in the ore-forming process causal knowledge graph comprises,

[0019] The ore-forming process of a typical ore deposit type is summarized by the large model, and an ore-forming process link is output;

[0020] An observable index set is established by generating different types of data observable indexes for each ore-forming node through the large model;

[0021] The multi-source geoscience data in the target mineral research area are mapped to a unified rasterized spatial representation, and a unified feature vector is obtained through feature splicing and dimension reduction methods;

[0022] The large model is used to give the relationship corresponding to the feature dimension for the observable index set of each ore-forming node, and a mapping relationship from the ore-forming node to the feature index set is established.

[0023] As a preferred scheme of the large model-based mineral resource prediction agent construction method of the present application, wherein the support and importance of the mineralization node in the mineralization process link include,

[0024] After receiving the mineral prediction request, based on the content of the mineral prediction request, the regional structure of the target mineral type and the target mineral research area is retrieved through the large model, and a set of mineralization process links of the target mineral research area is generated;

[0025] According to the feature mapping relationship established for each grid and each mineralization node, the support of the mineralization node in the grid observation in the mineralization process link is calculated through the feature sub-vector;

[0026] Based on the mineralization process causal knowledge graph, the out-degree and in-degree of each mineralization node in the mineralization process link is counted to obtain a structure score;

[0027] By counting the number of occurrences of the mineralization node in the mineralization process link in the set of mineralization process links, a frequency score is obtained;

[0028] Through the evaluation template of the large model, a semantic score of the mineralization node is output;

[0029] The structure score, frequency score and semantic score are unified as the importance of the mineralization node.

[0030] As a preferred scheme of the large model-based mineral resource prediction agent construction method of the present application, wherein the generated data gap index includes,

[0031] The data gap index calculation formula is constructed through the support and importance of the mineralization node;

[0032] The data gap index is calculated through the data gap index calculation formula.

[0033] As a preferred scheme of the large model-based mineral resource prediction agent construction method of the present application, wherein the generated active sampling task includes,

[0034] For the mineralization node with a data gap index greater than a data gap threshold, an observation index set is extracted;

[0035] The large model deduces a sampling type set according to the extracted observation index set;

[0036] The large model generates a sampling space position set according to the spatial data corresponding to the mineralization node and the grid features;

[0037] Each pair of sampling space position and sampling type set is constructed into an active sampling task set;

[0038] The active sampling task set includes a grid position and a sampling type corresponding to each active sampling task in the active sampling task set, and is associated with an ore-forming node and an ore-forming process link corresponding to the ore-forming node.

[0039] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the active sampling scheme includes,

[0040] An uncertainty measure is defined for each ore-forming node belonging to the ore-forming process link in the active sampling task set.

[0041] The uniform feature vector of the grid position corresponding to the active sampling task and the sampling type are observed by the mixture density network to obtain the observation value distribution of the active sampling task.

[0042] Based on the observation value distribution of the active sampling task, the posterior support of the ore-forming node on the ore-forming process link is obtained through Bayesian update.

[0043] According to the posterior support, the updated uncertainty measure is calculated.

[0044] According to the uncertainty measure of the ore-forming node in the ore-forming process link before and after the update, the expected information gain of the active sampling task is calculated.

[0045] Under the constraint condition of the given sampling budget, the active sampling task set with the highest expected information gain and meeting the cost constraint is selected as the execution scheme.

[0046] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the mixture density network includes,

[0047] For each type of sampling, a training sample is constructed for each historical observation to form a training data set.

[0048] The training sample includes a grid space position, a uniform feature vector, a sampling type, and an actual observation value.

[0049] The sampling type in the training data set is used as an embedding vector, and the uniform feature vector of the grid space position and the embedding vector are combined into a context vector.

[0050] The context vector is mapped to the parameter set of the mixture Gaussian through the neural network in the mixture density network, and the conditional distribution is defined based on the parameter set of the mixture Gaussian.

[0051] Based on the training process, the negative log-likelihood is used as the loss function of the mixture density network, and the parameters are trained by gradient descent method until the validation set converges.

[0052] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the Bayesian update comprises,

[0053] A binary random variable is defined in advance for the ore-forming node in the ore-forming process causal knowledge graph.

[0054] Based on the binary random variable, the support degree of the ore-forming node is taken as the prior probability.

[0055] For each ore-forming node, a set of random variable observation values is collected, and the set of random variable observation values is respectively fitted with a normal distribution.

[0056] When the observation value distribution of the active sampling task is obtained, the posterior support degree of the ore-forming node on the ore-forming process link is calculated according to the Bayesian formula.

[0057] The posterior support degree of the ore-forming node on the ore-forming process link is calculated to obtain the updated uncertainty measure of the ore-forming node in the ore-forming process link through the defined uncertainty measure.

[0058] As a preferred scheme of the large model-based mineral resource prediction agent construction method, the receiving of new geoscience data generated by active sampling, the updating of the unified feature vector, the recalculation of the data gap index under the control of the large model and the iterative updating are performed, and the specific steps are,

[0059] According to the sampling task in the active sampling scheme, new geoscience data is collected.

[0060] The geoscience data is converted into a feature form isomorphic to the unified feature vector to obtain the updated unified feature vector.

[0061] The support degree of the ore-forming node is updated using the updated unified feature vector.

[0062] Based on the accumulated geoscience data of multiple rounds of active sampling, the ore-forming process causal knowledge graph is updated.

[0063] Based on the updated ore-forming process causal knowledge graph and the support degree of the ore-forming node, the data gap index is recalculated, and new active sampling tasks and expected information gain are generated.

[0064] Under the conditions of resources and time, a next round of active sampling task combination is generated to form a cyclic iteration.

[0065] The present application has the beneficial effects that: by utilizing a large model to perform semantic analysis and relationship extraction on geological data and atlases of a target mineral species research area, a mineralization process causal knowledge graph is constructed, explicit representation and quantitative characterization of a mineralization event chain and key control nodes are realized; by introducing a mixed density network to predict the distribution of observation values of different sampling types, combining with Bayesian update of mineralization node uncertainty, expected information gain of an active sampling task is calculated, an active sampling scheme is optimized under sampling cost constraints, quantitative trade-off between exploration resource investment and information gain is realized, and the comprehensive efficiency of the sampling scheme in terms of cost and benefit is improved. BRIEF DESCRIPTION OF DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0067] Fig. 1 The flowchart of the method for constructing a mineral resource prediction agent based on a large model.

[0068] Fig. 2 The flowchart of constructing a mineralization process causal knowledge graph.

[0069] Fig. 3 The flowchart of establishing a unified feature vector corresponding to a mineralization node.

[0070] Fig. 4 The flowchart of geoscience data iterative update. DETAILED DESCRIPTION

[0071] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.

[0072] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0073] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.

[0074] ReferenceFigs. 1-4 For an embodiment of the present application, the embodiment provides a large model-based mineral resource prediction agent construction method, comprising the following steps:

[0075] S1, obtain the geological data and atlas of the target mineral research area, use the large model to extract the text structure of the mineralization process, construct the causal knowledge graph of the mineralization process, and rasterize and feature encode the multi-source geoscience data to establish a unified feature vector corresponding to the mineralization node in the causal knowledge graph of the mineralization process.

[0076] Further, collect data sources to obtain the geological data and atlas of the target mineral research area.

[0077] The data sources include geological reports of the target mineral research area, geological maps of the target mineral research area, and international and domestic summary data of mineralization regularities of the target mineral research area.

[0078] Among them, the geological reports of the target mineral research area include but are not limited to mineral exploration reports, mineral deposit research papers, and mineral deposit monographs; the geological maps of the target mineral research area include but are not limited to geological structure maps, magmatic rock distribution maps, mineralization belt division maps, and accompanying explanatory texts; the international and domestic summary data of mineralization regularities of the target mineral research area include but are not limited to mineralization mode maps, mineralization series division, and typical mineral deposit mineralization modes.

[0079] Further, a pre-trained large model is selected as a basic model, and the geological data and atlas of the target mineral research area are cleaned, de-duplicated, and formatted. The large model is trained in the field to enable the large model to understand and generate sentences and structured outputs with geoscience semantics.

[0080] Through targeted prompt templates, the large model is guided to structure the text related to the mineralization process to obtain mineralization nodes, mineralization causal relationships, and mineralization time sequence relationships.

[0081] Among them, the mineralization nodes include event class nodes and material class nodes; the event class nodes include but are not limited to tectonic events, magmatic events, metamorphic events, hydrothermal activities, and sedimentary events; the material class nodes include but are not limited to specific lithology, mineral assemblage, alteration type, and fluid characteristics.

[0082] The mineralization time sequence relationship includes but is not limited to the order of sequence, coexistence in the same period, and superimposed modification.

[0083] The mineralization causal relationship includes but is not limited to tectonic extension leading to the upward movement of ore-bearing magma and acidic rock body intrusion leading to contact metasomatism mineralization.

[0084] It should be noted that the construction process of the prompt template is determined, specifically, the extraction tasks of metallogenic event extraction, metallogenic causal relationship extraction, metallogenic time sequence relationship extraction and observation index extraction are determined respectively corresponding task identification and field structure; wherein the field structure at least includes event name field, relationship type field, time sequence relationship field and text segment field, to form the extraction slot for constraining the output format of the large model. Select a plurality of example text segments capable of representing different metallogenic event relationships from pre-collected geological data and album data, one-to-one correspondence between each example text segment and the corresponding field structure, generate example input and example output pairs, and construct a prompt template preliminary draft containing task description, output format constraint and a number of example input and output pairs on the basis of example input and example output pairs; based on the pre-constructed small-scale labeled geological corpus, calling the large model and loading the prompt template preliminary draft to extract metallogenic events and metallogenic relationships from the labeled corpus, comparing the large model output with the labeled results, calculating the field completeness rate and the extraction accuracy, modifying and supplementing the task description, output format constraint and example input and output pairs in the prompt template preliminary draft according to the comparison results, until the extraction accuracy and the field completeness rate meet the requirements, and the target prompt template is obtained.

[0085] The structured extraction result is stored to form a directed graph structure of the metallogenic process causal knowledge graph, which is represented as:

[0086]

[0087] Among them, is the metallogenic process causal knowledge graph, is the metallogenic node set, represents the directed edge set.

[0088] The metallogenic node set is used to represent the metallogenic event, and the directed edge set is used to represent the metallogenic causal relationship; wherein each directed edge can be accompanied by a [0, 1] confidence parameter to reflect the literature frequency and expert review results.

[0089] Further, after obtaining the metallogenic process causal knowledge graph, the large model is used to summarize the metallogenic process of typical ore deposit types, specifically, for each ore deposit type, such as porphyry copper deposit, carlin type gold deposit, VMS deposit, etc., through the metallogenic chain template, the large model outputs the common metallogenic step sequence of the ore deposit, and then maps it to the metallogenic process link, which is represented as:

[0090]

[0091] Among them, represents the th metallogenic process link of the th ore deposit, represents the ​​Indexing the ore-forming nodes on the ore-forming process chain link, representing the first Ore-forming nodes of the ore-forming process chain.

[0092] For each ore-forming node, generate observable indicators in different types of data through a large model, such as corresponding lithology, structural form which can be observed in geological map and drilling lithology; corresponding physical property difference which can be observed in magnetic, electric, gravity data; corresponding alteration and mineral combination which can be observed in spectral remote sensing.

[0093] Establish a set of observation indicators for each ore-forming node.

[0094] Further, map the existing multi-source geoscience data in the target mineral species research area to a unified rasterized spatial representation, so as to calculate the ore-forming causal support degree of each spatial position.

[0095] Specifically, divide the target mineral species research area into a two-dimensional grid set, and each grid corresponds to a fixed plane position.

[0096] Extract corresponding features from different data sources for each grid, including but not limited to: geological feature vector, geophysical feature vector, geochemical feature vector, remote sensing feature vector, and other supplementary features.

[0097] Among them, the geological feature vector includes but is not limited to stratum, lithology, tectonic unit, intrusive body contact relationship, etc. The geophysical feature vector includes but is not limited to gravity, magnetic method, electric method, seismic attribute, etc. after standardization. The geochemical feature vector includes but is not limited to element concentration, ratio, primary and secondary element combination, etc. The remote sensing feature vector includes but is not limited to multispectral / hyperspectral band combination, alteration information, etc. Other supplementary features include but are not limited to known mine distance, fault distance, topographic attribute, etc.

[0098] Get a unified feature vector through feature splicing and dimension reduction methods such as principal component analysis.

[0099] Use a large model to give a corresponding relationship between the observation indicator set of each ore-forming node and the feature dimension.

[0100] For example, the large model outputs a corresponding feature index set according to the prompt "Map the 'high magnetic anomaly' indicator to which dimensions in the geophysical feature".

[0101] Based on the output of the large model, establish a mapping from the ore-forming node to the feature index set.

[0102] S2, under the control of the large model, based on the ore-forming process causal knowledge graph and the unified feature vector, reason the ore-forming process chain link, and based on the ore-forming process chain link, calculate the support degree and importance of the ore-forming nodes in the ore-forming process chain link.

[0103] Further, after receiving the mineral prediction request, the request content is first analyzed by the large model, specifically, the spatial range, geographic location, regional geological background sketch, target mineral type, constraint conditions, etc. of the target mineral research area are input into the large model.

[0104] The large model retrieves and generates a set of possible ore-forming process links for the target mineral research area according to the target mineral type and the regional structure of the target mineral research area. Each link is a directed path from the ore-forming source to the mineralization precipitation and the later modification.

[0105] Further, the process of generating candidate links is specifically implemented as steps A1-A4:

[0106] A1: The large model calls a pre-constructed set of ore-forming process links for the target mineral.

[0107] A2: According to the tectonic system, magmatic series and regional metamorphic grade information of the target mineral research area, the set of ore-forming process links is screened and adjusted.

[0108] A3: The ore-forming process links that do not conform to the regional geological evolution are marked as not applicable by the large model and removed.

[0109] A4: Output a set of ore-forming process links suitable for the study area.

[0110] Further, for each grid and each ore-forming node, the observation support degree of the ore-forming node in the grid is calculated according to the established feature mapping relationship.

[0111] Specifically, for each type of ore-forming node, the corresponding recognition model is called, and for lithology and structure features, a graph convolution-based classifier is used; for geophysical anomaly patterns, a two-dimensional convolutional network is used; for geochemical composite anomalies, a deep density estimation model is used; for remote sensing alteration information, a convolutional network is used.

[0112] It should be noted that the training process of each recognition model is as follows: for lithology and structure features, a spatial adjacency graph is constructed based on the geological grid, multi-source geoscience features are used as node features, and known lithology and structure types are used as supervision labels, and a graph convolutional neural network is used for supervised training; for geophysical anomaly patterns, two-dimensional image blocks containing anomalies and backgrounds are cropped from geophysical grid data, and anomaly types are used as labels, and a two-dimensional convolutional neural network is used for supervised training; for geochemical composite anomalies, a variational autoencoder and other deep density estimation models are trained based on background samples far from known ore bodies, and reconstruction error is used as an anomaly degree indicator; for remote sensing alteration information, image blocks with alteration type labels are cropped from remote sensing images, and a convolutional neural network is used for supervised training, so that each recognition model meets the accuracy on the validation set, and is used as a generation model for the observation support degree of the ore-forming node.

[0113] The feature sub-vector corresponding to the grid is input into the recognition model to obtain a matching probability of the ore-forming node in the grid, and the matching probability of the ore-forming node in the grid is taken as the support degree of the ore-forming node.

[0114] For the importance of the ore-forming node of each ore-forming process link, in the ore-forming process causal knowledge graph, the number of causal edges pointing to the ore-forming node and the number of causal edges from the ore-forming node to the ore-forming event are counted, the number of causal edges pointing to the ore-forming node is taken as the in-degree, the number of causal edges from the ore-forming node to the ore-forming event is taken as the out-degree, and the normalized value of the sum of the in-degree and the out-degree is taken as the structure score.

[0115] The number of occurrences of the ore-forming node in the ore-forming process link in the ore-forming process link set is counted, and the ratio of the number of occurrences in the ore-forming process link to the total number of ore-forming process links is taken as the frequency score.

[0116] For the semantic score, a fixed evaluation template is given to the large model, the large model returns structured output according to the evaluation standard in the evaluation template, the structured output output by the large model is linearly mapped to obtain the semantic score.

[0117] The evaluation template can be "From the perspective of controlling mineral precipitation or metal enrichment, evaluate the importance of the ore-forming node The importance of the corresponding geological event or geological body in the ore-forming process of the target ore is given as an integer in the 0~4 point system, where 0 represents almost nothing, 1 represents secondary, 2 represents general importance, 3 represents a key factor, and 4 represents a core control factor, and a brief reason is given."

[0118] Finally, the structure score, the frequency score and the semantic score are summed by weighting to unify the importance index; wherein the structure score weight, the frequency score weight and the semantic score weight are all non-negative weights and the sum is 1, which is obtained by optimizing the prediction effect and then obtaining the importance index.

[0119] S3, according to the support degree and importance of the ore-forming node, generate a data gap index, derive the sampling type and spatial position needed to fill the data gap by the large model, generate an active sampling task, and optimize it combined with the sampling cost and expected information gain to obtain an active sampling scheme.

[0120] Further, the data gap index is constructed by the support degree and importance of the ore-forming node, and is represented as:

[0121] ;

[0122] Wherein, represents the importance of the ore-forming node a data gap indicator, an importance of the ore-forming node , a support of the ore-forming node .

[0123] Further, for the ore-forming node with the data gap indicator greater than the data gap threshold, an observation indicator set is extracted, and the large model derives a possible sampling type set according to the observation indicator set of the ore-forming node.

[0124] For example, for the ore-forming node with unclear deep intrusive rock body morphology, the derived sampling types include new gravity and magnetic survey lines and three-dimensional magnetic inversion; for the ore-forming node lacking alteration mineral combination constraints, the derived sampling types include remote sensing encryption interpretation, surface intensive sampling, and rock and mineral experimental analysis.

[0125] It should be noted that the data gap threshold can be a preset threshold, and the data gap threshold value range is generally [0, 1], or the upper quartile point of the data gap indicator distribution obtained by statistical analysis on the historical deposit data region can be taken as the data gap threshold.

[0126] The large model generates a sampling spatial position set based on existing spatial data and grid features.

[0127] For example, for geophysical sampling tasks, equidistant sampling can be generated according to geological boundaries, existing survey line spacing, and terrain within the coverage area; for drilling sampling tasks, candidate drilling locations can be generated around the high ore-forming probability area location; for surface sampling tasks, candidate points can be generated in the grid lacking geochemical exploration points.

[0128] Each pair of sampling spatial position and sampling type set is constructed as an active sampling task set; each task in the active sampling task set corresponds to a grid position and a sampling type, and is associated with an ore-forming node and an ore-forming process link.

[0129] Further, the expected information gain of the active sampling task set is estimated, and the active sampling task set is sorted based on the expected information gain.

[0130] Specifically, an uncertainty metric is defined for each ore-forming node on the ore-forming process link, which is represented as:

[0131] ;

[0132] Wherein, represents the th ore-forming process link, represents the overall uncertainty indicator of the ore-forming process link , and represents the ore-forming node index on the th ore-forming process link, Represents the mineralization process chain The first One mineralized node Representing mineralization nodes The importance of Representing mineralization nodes Support It is an uncertain function.

[0133] The uncertainty function can be expressed as:

[0134] ;

[0135] Furthermore, in order to estimate the information gain after the sampling task is performed, it is necessary to predict the possible observations of the current sampling task.

[0136] In this embodiment, a hybrid density network is used to predict the possible observations for the current sampling task. Specifically, the hybrid density network is constructed by creating a training sample for each historical observation of each sampling type. The training sample includes the grid spatial location, a unified feature vector, the sampling type, and the actual observation, forming a training dataset. The sampling type in the training dataset is used as an embedding vector, and the unified feature vector of the grid spatial location and the embedding vector are combined into a context vector. Let the observation value... The dimension is Through neural networks in hybrid density networks The context vector is mapped to the parameter set of the Gaussian mixture, and a conditional distribution is defined based on the parameter set of the Gaussian mixture. The negative log-likelihood is used as the loss function of the mixture density network during the training process, and the parameters are trained by gradient descent until the validation set converges.

[0137] It should be noted that the parameter set of the Gaussian mixture is expressed as:

[0138] ;

[0139] in, Indicates parameters Hybrid density network, Represents the context vector. This represents the output obtained by inputting the context vector into the hybrid density network. Indicates the first The weights of the Gaussian components, with values ​​ranging from [0,1]. Indicates the first The mean vector of Gaussian components, Indicates the first The covariance matrix of Gaussian components, denotes the number of selected Gaussian components in the mixture density network.

[0140] The conditional distribution is denoted as:

[0141] ;

[0142] wherein, denotes the unified feature vector and the embedding vector under the condition of the observation value , the conditional prediction probability of denotes the unified feature vector corresponding to the grid space position, denotes the embedding vector, denotes the density value of the multivariate normal distribution with as the mean and as the covariance at the point .

[0143] The loss function of the mixture density network is denoted as:

[0144] ;

[0145] wherein, denotes the loss function used during training of the mixture density network, denotes the total number of training samples, denotes the position unified feature vector corresponding to the th training sample, denotes the embedding vector corresponding to the th training sample, denotes the observation value corresponding to the th training sample, denotes the conditional probability density predicted by the mixture density network under the parameters for the sample , i.e., the value at the observation value .

[0146] Further, based on the obtained active sampling task, the unified feature vector is extracted at the grid, the sampling type is converted into an embedding vector, the unified feature vector and the embedding vector are spliced into a context vector, and the context vector is input into the mixture density network to output the observation value distribution of the active sampling task.

[0147] After obtaining the distribution of observations from the active sampling task, the uncertainty of the ore-forming nodes in the mineralization process chain is updated using Bayesian updates. Specifically, for each ore-forming node in the causal knowledge graph of the mineralization process, a binary random variable is predefined. A random variable of 1 indicates that the mineralization event corresponding to the ore-forming node is true in the current study area, and a random variable of 0 indicates that the mineralization event corresponding to the ore-forming node is not true. The support of the ore-forming node is used as the prior probability. For each ore-forming node, a set of observations with a random variable of 1 and a set of observations with a random variable of 0 are collected, and normal distributions are fitted to the two sets of observations respectively. When the distribution of observations from the active sampling task is obtained, the posterior support of the ore-forming nodes in the mineralization process chain is calculated according to the Bayesian formula, expressed as:

[0148] ;

[0149] ;

[0150] ;

[0151] ;

[0152] in, Representing mineralization nodes The corresponding binary state variable, This indicates that the mineralization event has been established. This indicates that the mineralization event is not valid. This represents the observations acquired by the current active sampling task. This represents the posterior support after sampling in the active sampling task. This represents the prior support before sampling in an active sampling task. Indicates the hypothetical mineralization node The probability density under the condition that it holds true. Indicates the hypothetical mineralization node The probability density under the condition that the condition is not true. Indicates in The mean of the observed values ​​under the given conditions Indicates in The standard deviation of the observed values ​​under these conditions Indicates in The mean of the observed values ​​under the given conditions Indicates in Standard deviation of observed values ​​under certain conditions It is an abbreviation for probability.

[0153] It should be noted that Bayesian updates use scalar statistics extracted from the observations.

[0154] The posterior support degree of the mineralization node in the mineralization process link is substituted into the mineralization process link uncertainty measurement formula to obtain the updated uncertainty measurement of the mineralization node in the mineralization process link.

[0155] The expected information gain of the active sampling task is calculated according to the uncertainty measurements of the mineralization nodes in the mineralization process link before and after the update, and the expected information gain is represented as:

[0156] ;

[0157] wherein, the expected information gain of the active sampling task is represented as: the i-th active sampling task, is represented as: is a mathematical expectation operator, and represents an expected value under the condition that the observation value obeys the distribution is represented as: the observation value obeys the probability distribution, is represented as: is represented as: is represented as: is represented as: is represented as: is represented as: is represented as: is represented as: is represented as: is represented as:

[0158] wherein, the updated uncertainty measurement is represented as:

[0159] ;

[0160] wherein, is represented as: is represented as:

[0161] Further, after estimating the expected information gain of each active sampling task, the cost of performing the active sampling task is taken as a constraint condition, and the active sampling task set with the highest expected information gain is selected from the active sampling tasks satisfying the constraint condition as an execution scheme.

[0162] ​​S4, receiving new geoscience data generated by active sampling, updating the unified feature vector, re-computing the data gap index under the control of the large model and iteratively updating.

[0163] Further, after the completion of the active sampling task, the new geoscience data obtained is converted into a feature form isomorphic to the original unified feature vector, and an updated grid unified feature vector is obtained.

[0164] The support of the ore-forming node is re-computed using the updated unified feature vector.

[0165] After multiple rounds of active sampling, the large model locally corrects and extends the original ore-forming process causal knowledge graph based on the accumulated new geoscience data, specifically, gradually reduces the confidence of the ore-forming relationship edges that have been repeatedly falsified or weakened; through the analysis report and data results of the large model, the corresponding ore-forming nodes, ore-forming causal relationships and ore-forming temporal relationships are extracted and added for new ore-forming relationships that repeatedly appear in multiple regions and projects but are not explicitly expressed in the initial knowledge base.

[0166] After updating the ore-forming causal structure and the support of the ore-forming node, the data gap index of each grid in the study area is re-computed based on the updated support of the ore-forming node, the active sampling task and the expected information gain are re-generated, and the next round of active sampling task combination is generated under the condition of resource and time allowing, forming a cycle iteration.

[0167] In summary, the present application realizes the explicit representation and quantitative characterization of the ore-forming event chain and key control nodes by using the large model to perform semantic analysis and relationship extraction on the geological data and atlas of the target mineral research area, and constructing the ore-forming process causal knowledge graph; the present application realizes the fusion expression of heterogeneous data such as geology, geophysical prospecting, geochemical prospecting and remote sensing in a unified spatial framework, improves the comprehensive utilization degree and spatial quantitative analysis capability of ore-forming information, and improves the comprehensive efficiency of sampling schemes in terms of cost and benefit by establishing the unified feature vector corresponding to the ore-forming node based on the unified feature vector, computing the support and importance of the ore-forming node, constructing the data gap index, deriving the sampling type and spatial location under the control of the large model, generating the active sampling task corresponding to the ore-forming process link, calculating the expected information gain of the active sampling task by introducing the mixed density network to predict the observation value distribution of different sampling types and combining the Bayesian update of the uncertainty of the ore-forming node, and optimizing and selecting the active sampling scheme under the constraint of sampling cost.

[0168] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A method for constructing a mineral resource prediction intelligent agent based on a large model, characterized in that: include, Geological data and maps of the target mineral study area are obtained, and the text structure of the mineralization process is extracted using a large model to construct a causal knowledge graph of the mineralization process. Multi-source geological data are rasterized and feature-encoded to establish a unified feature vector corresponding to the mineralization nodes in the causal knowledge graph of the mineralization process. Under the control of a large model, based on the causal knowledge graph of the mineralization process and the unified feature vector, the mineralization process chain is inferred, and based on the mineralization process chain, the support and importance of the mineralization nodes in the mineralization process chain are calculated; Based on the support and importance of mineralized nodes, a data gap index is generated. The sampling type and spatial location required to fill the data gap are derived from the large model, an active sampling task is generated, and the active sampling scheme is obtained by combining the sampling cost and expected information gain. It receives new geoscientific data generated by active sampling, updates the unified feature vector, and recalculates the data gap index under the control of the large model and performs iterative updates. The support and importance of ore-forming nodes in the calculation of the ore-forming process chain include, Upon receiving a mineral prediction request, based on the content of the mineral prediction request, a large model is used to search for the target mineral type and the regional tectonics of the target mineral study area, and a set of metallogenic process links in the target mineral study area is generated. Based on the feature mapping relationship established for each grid and each ore-forming node, the support of the ore-forming node in the ore-forming process link in the grid observation is calculated by the feature sub-vectors. Based on the causal knowledge graph of the mineralization process, the in-degree statistics of each mineralization node in the mineralization process link are performed to obtain a structural score. Frequency scores are obtained by counting the number of times each ore-forming node appears in the ore-forming process chain in the ore-forming process chain set. Semantic scores are output for ore-forming nodes using the evaluation template of the large model. The structural score, frequency score, and semantic score are unified to determine the importance of ore-forming nodes.

2. The method for constructing a mineral resource prediction agent based on a large model as described in claim 1, characterized in that: The constructed causal knowledge graph of the mineralization process includes... Collect data sources to obtain geological data and maps of the target mineral study area; The geological data and maps of the target mineral study area were cleaned, deduplicated, and formatted. The large model was adapted and trained based on the processed geological data and maps of the target mineral study area. The text of the mineralization process is input into a large model that has been adapted and trained. The structure is extracted using prompt templates to obtain mineralization nodes, causal relationships of mineralization, and temporal relationships of mineralization. The mineralization nodes, causal relationships, and temporal relationships are stored in a directed graph structure to construct a causal knowledge graph of the mineralization process. In the causal knowledge graph of the mineralization process, mineralization nodes are used to represent mineralization events, and directed edges are used to represent causal relationships in mineralization.

3. The method for constructing a mineral resource prediction agent based on a large model as described in claim 2, characterized in that: The unified feature vectors corresponding to the ore-forming nodes in the causal knowledge graph of the ore-forming process include... The mineralization process of typical ore deposit types is summarized by large-scale models, and the mineralization process chain is output. For each mineralization node, an observable index of different types of data is generated through a large model, and a set of observable indices is established. Multi-source geoscientific data within the target mineral study area are mapped into a unified rasterized spatial representation, and a unified feature vector is obtained through feature splicing and dimensionality reduction methods. By using a large model, the relationship between the set of observation indicators for each mineralization node and the feature dimension is given, and a mapping relationship from mineralization node to feature index set is established.

4. The method for constructing a mineral resource prediction agent based on a large model as described in claim 3, characterized in that: The generated data gap indicators include, A formula for calculating the data gap index is constructed by considering the support and importance of ore-forming nodes; The data gap index is calculated using the formula for calculating the data gap index.

5. The method for constructing a mineral resource prediction agent based on a large model as described in claim 4, characterized in that: The generation of the active sampling task includes, For mineralized nodes where the data gap index is greater than the data gap threshold, extract the set of observation indices. The sampling type set is derived by using a large model based on the extracted set of observation indicators. A set of sampled spatial locations is generated using a large model based on the spatial data and raster features corresponding to the ore-forming nodes; For each pair of sampling spatial locations and sampling types, an active sampling task set is constructed; The active sampling task set includes each active sampling task in the active sampling task set corresponding to a grid position and a sampling type, and is associated with a mineralization node and the mineralization process link corresponding to the mineralization node.

6. The method for constructing a mineral resource prediction agent based on a large model as described in claim 5, characterized in that: The active sampling acquisition scheme includes, Define an uncertainty measure for each mineralization node in the active sampling task set; For the unified feature vector and sampling type of the grid position corresponding to the active sampling task, the observation distribution of the active sampling task is obtained by predicting the observation value through a hybrid density network. The distribution of observations based on the active sampling task is updated using Bayesian methods to obtain the posterior support of ore-forming nodes in the ore-forming process chain. Calculate the updated uncertainty measure based on the posterior support; The expected information gain of the active sampling task is calculated based on the uncertainty measure before and after the update of the ore-forming nodes in the ore-forming process chain. Given a sampling budget constraint, the set of active sampling tasks that has the highest expected information gain and meets the cost constraint is selected as the execution plan.

7. The method for constructing a mineral resource prediction agent based on a large model as described in claim 6, characterized in that: The hybrid density network includes, For each historical observation of each sampling type, a training sample is constructed to form a training dataset; The training samples include grid spatial location, uniform feature vector, sampling type, and actual observation values; The sampling type in the training dataset is used as the embedding vector, and the unified feature vector of the grid spatial location and the embedding vector are combined into a context vector. The context vector is mapped to the parameter set of the Gaussian mixture through the neural network in the mixture density network, and the conditional distribution is defined based on the parameter set of the Gaussian mixture. The training process uses negative log-likelihood as the loss function for the hybrid density network, and the parameters are trained using gradient descent until the validation set converges.

8. The method for constructing a mineral resource prediction agent based on a large model as described in claim 7, characterized in that: The Bayesian update includes, For the ore-forming nodes in the causal knowledge graph of the ore-forming process, a binary random variable is predefined; Based on binary random variables, the support of ore-forming nodes is used as the prior probability; For each mineralized node, a set of random variable observations is collected, and a normal distribution is fitted to the set of random variable observations. When obtaining the distribution of observations for an active sampling task, the posterior support of the ore-forming nodes in the ore-forming process link is calculated according to Bayes' theorem. The posterior support of the ore-forming nodes in the ore-forming process chain is calculated using a defined uncertainty metric to obtain the updated uncertainty metric of the ore-forming nodes in the ore-forming process chain.

9. The method for constructing a mineral resource prediction agent based on a large model as described in claim 8, characterized in that: The specific steps are as follows: receiving new geoscientific data generated by active sampling, updating the unified feature vector, recalculating the data gap index under the control of the large model, and iteratively updating it. Collect new geoscientific data according to the sampling tasks in the active sampling plan; Transform geoscientific data into a feature form isomorphic to a unified feature vector to obtain the updated unified feature vector; The support of ore-forming nodes is updated using the updated unified feature vector; Based on the accumulated geoscientific data from multiple rounds of active sampling, the causal knowledge graph of the mineralization process is updated; Based on the updated causal knowledge graph of the mineralization process and the support of mineralization nodes, the data gap index is recalculated, and new active sampling tasks and expected information gains are regenerated. If resources and time permit, generate the next round of active sampling task combinations to form a loop iteration.

Citation Information

Patent Citations

  • Protection method and device for underground cable

    CN118521042A

  • Hidden ore body evaluating and positioning method based on multi-source data processing

    CN121094349A