Geologic constraint-based graph reinforcement learning mineral product prediction method

By employing a geologically constrained graph reinforcement learning method, utilizing geometric graph structures and graph convolutional neural networks, and combining known mineralization information and geological knowledge, the interaction between the agent and the environment is optimized. This solves the problem of insufficient integration of multi-source information in mineral prediction and achieves higher recognition accuracy.

CN120805973APending Publication Date: 2025-10-17CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510630740.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing deep learning methods for mineral prediction are insufficient in integrating multi-source geological information and lack deep interaction between various geological prospecting information and geological knowledge, resulting in low accuracy in mineral prediction and identification.

Method used

We employ a geologically constrained graph reinforcement learning method. By constructing a geometric graph structure as the state, we combine reward signals from known mineralization information, potential mineralization information, and geological knowledge. We utilize graph convolutional neural networks and Markov chain decision processes to optimize the interaction between the agent and the environment. We design a loss function to train the model and improve the accuracy of mineral prediction.

Benefits of technology

It effectively improves the identification accuracy of mineral resource prediction, can better integrate multi-source geological information, conforms to the metallogenic regularity, and improves the accuracy of mineral resource prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805973A_ABST
    Figure CN120805973A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of earth science, and particularly discloses a geological constraint-based graph reinforcement learning mineral product prediction method, which comprises the following steps of: constructing states, rewards and actions in graph reinforcement learning, adding geological constraints in a process of constructing graph reinforcement learning environment reward feedback, rewards are fed back to the intelligent agent by the environment under the constraint of geological knowledge, and the intelligent agent is promoted to make a multi-angle decision on the mineralization potential; constructing a Markov chain decision process of graph reinforcement learning, and realizing deep coupling of an intelligent agent and an environment; the method comprises the following steps: designing a loss function, establishing a target for learning of an intelligent agent, training a model by continuously reducing a gap between a target network and an intelligent agent network, gradually identifying and distinguishing known mineralization, potential mineralization and non-mineralization in an interaction process of the intelligent agent and an environment, and delineating a metallogenic prospective area. The method can effectively improve the recognition precision of model mineral prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of earth science, and more particularly, relates to a geological constraint-based graph reinforcement learning mineral prediction method. BACKGROUND

[0002] Mineral prediction work is based on the research on ore-forming dynamics and ore-forming regularity, integrates multi-source geological prospecting data such as geology, geophysics, geochemistry, remote sensing, drilling, extracts and fuses ore-forming and ore-indicating information, and delineates prospecting areas. In recent years, with the development of big data and deep learning, deep learning has successfully applied to mineral prediction due to its advantage of automatically extracting high-level and abstract mineralization features. According to different learning methods, the currently mainstream deep learning methods can be divided into supervised learning, unsupervised learning and semi-supervised learning.

[0003] In supervised learning, the model is trained on labeled data to learn the mapping relationship between input and output, thereby realizing prediction. Supervised learning focuses on learning the information of labels and is prone to ignore potential geological anomalies. Unsupervised learning does not rely on labeled data and mainly learns according to the internal structure of data, aiming to discover the potential patterns, relationships or structures of data, but often fails to fully utilize existing mineralization information or prior mineralization knowledge. Semi-supervised learning is a combination of supervised learning and unsupervised learning, and is usually used in the case of data label scarcity or high labeling cost. Through the joint learning of a small amount of labeled data and a large amount of unlabeled data, the generalization ability of the model is further improved. Although semi-supervised methods combine supervised learning and unsupervised learning, they still mainly focus on the features of labeled data and are prone to ignore the potential anomaly information contained in the data itself.

[0004] Overall, the current deep learning method for mineral prediction still has deficiencies in integrating multi-source geological information, and lacks deep interaction between multi-aspect geological prospecting information and geological knowledge. At present, how to construct a deep learning model that can simultaneously integrate known mineralization information, potential mineralization information and geological knowledge, and thereby improve the accuracy of mineral prediction, is a problem that needs to be studied. SUMMARY

[0005] In view of the defects of the prior art, the purpose of the present application is to provide a geological constraint-based graph reinforcement learning mineral prediction method, which can effectively improve the recognition accuracy of model mineral prediction.

[0006] To achieve the above purpose, the present application provides a geological constraint-based graph reinforcement learning mineral prediction method, comprising the following steps:

[0007] S10, constructing the state, reward and action in graph reinforcement learning, and the construction method is:

[0008] S11, a geological prospecting dataset is constructed according to a metallogenic system, and a state of a graph reinforcement learning environment is constructed from the geological prospecting dataset;

[0009] S12, a reward feedback of the graph reinforcement learning is constructed, and a geological constraint is added in the construction process to realize that the environment feeds back a reward to the agent under the constraint of geological knowledge;

[0010] S13, an agent of the graph reinforcement learning is constructed, and the agent and the environment are intelligently interacted, and whether the state has metallogenic potential is determined according to the state and the reward fed back by the environment, so as to fit a nonlinear relationship between the geological prospecting data and the metallogenic potential;

[0011] S20, a Markov chain decision process of the graph reinforcement learning is constructed to realize deep coupling of the agent and the environment;

[0012] S30, a loss function is designed to establish a target for learning of the agent, the model is trained by continuously reducing a gap between a target network and a network of the agent, and known mineralization, potential mineralization and non-mineralization are gradually identified and distinguished in the interaction process of the agent and the environment to delineate a metallogenic prospective area.

[0013] The method for mineral prediction based on the graph reinforcement learning under the geological constraint provided in the application uses a geometric graph to represent prospecting data and a complex spatial coupling relationship between the prospecting data and mineralization, and considers metallogenic potential of a complex environment from the aspects of known metallogenic information, potential metallogenic information and geological knowledge through interaction and coupling among states, actions and rewards of the reinforcement learning, so that the recognition accuracy of model mineral prediction can be effectively improved.

[0014] As a further optimization, in step S11, the geological prospecting dataset is rasterized, and on this basis, the raster data is converted into a geometric graph structure as a state of the graph reinforcement learning according to a set distance threshold.

[0015] As a further optimization, the geological prospecting dataset is rasterized by using GIS technology.

[0016] As a further optimization, in step S12, the known metallogenic information, the potential metallogenic information and the geological knowledge are taken as reward signals and are deeply coupled to be taken as a standard for the agent to determine whether a state has metallogenic potential.

[0017] As a further optimization, the geological knowledge includes metallogenic regularity.

[0018] As a further optimization, the agent adopts a graph convolutional neural network.

[0019] As a further preferred, in step S20, the interaction process between the agent and the environment is modeled as a Markov decision process, and a state transition function is established, so that the agent continuously adjusts by trial and error, and finally seeks the optimal mineralization potential prediction strategy.

[0020] As a further preferred, in step S30, the loss function L is:

[0021] L=E[(R t +γmaxQ′(s t+1 ,a t+1 )-Q(s t ,a t )) 2 ]

[0022] In the formula, t and t+1 represent the time sequence on the Markov chain; E represents the mean value; R t represents the reward at time t; s t+1 represents the state at time t+1; a t+1 represents the action at time t+1; s t represents the state at time t; a t represents the action at time t; and γ represents a discount factor, which balances the relationship between the current reward and the future reward; Q represents the Q value fitted by the agent network; and Q' represents the Q value fitted by the target network. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a flowchart of a mineral prediction method based on graph reinforcement learning under geological constraints provided by the embodiments of the present application;

[0024] Figure 2 is a structural diagram of a mineral prediction method based on graph reinforcement learning under geological constraints provided by the embodiments of the present application. DETAILED DESCRIPTION

[0025] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0026] In view of the problems that the traditional mineral prediction deep learning method is difficult to fully integrate multi-source geological prospecting information under complex geological environment, and the recognition accuracy is not high, the present application provides a mineral prediction method based on graph reinforcement learning under geological constraints, which uses a geometric graph to represent the prospecting data and the complex spatial coupling relationship between the prospecting data and mineralization, and comprehensively considers the mineralization potential of the complex environment from the perspective of known mineralization information, potential mineralization information and geological knowledge through the interactive coupling between the state, action and reward of reinforcement learning, so as to improve the recognition accuracy of the model mineral prediction.

[0027] As Figure 1 shown, the geological constraint-based graph reinforcement learning mineral prediction method provided in the application includes steps S10-S30, which are described in detail as follows:

[0028] Step S10, constructing the state, reward and action in the graph reinforcement learning.

[0029] It should be noted that in the graph reinforcement learning, the environment, agent, state, action and reward are the core components, among which the environment and agent interact continuously to realize the coupling of the state, action and reward. The construction sequence of the state, reward and action in the graph reinforcement learning provided in the embodiment is not limited.

[0030] In step S10, the construction method of the state, reward and action in the graph reinforcement learning is specifically:

[0031] Step S11, constructing a geological prospecting data set according to a mineralization system, and establishing the state of the graph reinforcement learning environment therefrom.

[0032] In the embodiment, the geological prospecting data set can be rasterized, and on this basis, the raster data is converted into a geometric graph structure as the state of the graph reinforcement learning model according to a set distance threshold. Specifically, the geological prospecting data set can be rasterized by using GIS technology.

[0033] Step S12, constructing the reward feedback of the graph reinforcement learning, in which the geological constraint is added to realize the feedback of the environment to the agent under the constraint of the geological knowledge, and to promote the agent to make multi-angle decisions on the mineralization potential.

[0034] In the embodiment, the known mineralization information, potential mineralization information and geological knowledge can be constructed as reward signals and deeply coupled to serve as the standard for the agent to judge whether the state has mineralization potential. Specifically, the geological knowledge includes the mineralization regularity.

[0035] Step S13, constructing the agent of the graph reinforcement learning, which intelligently interacts with the environment, executes action decision to judge whether it has mineralization potential according to the state and reward feedback from the environment, so as to fit the nonlinear relationship between the geological prospecting data and the mineralization potential.

[0036] In the embodiment, the agent of the graph reinforcement learning is constructed so that the agent can directly learn the relationship between the geometric graph data and the action (judging whether it has mineralization potential) and the reward, thereby fully learning the nonlinear relationship between the geological prospecting data and the mineralization potential. The agent can be modeled as a graph convolutional neural network for capturing the coupling relationship between the nodes and edges in the geometric graph.

[0037] Step S20, a Markov chain decision process of map reinforcement learning is constructed to realize deep coupling of the agent and the environment.

[0038] In step S20, the interaction process of the agent and the environment can be modeled as a Markov decision process, and a state transition function is established, so that the agent can continuously adjust through trial and error, and finally seek the optimal metallogenic potential prediction strategy.

[0039] Step S30, a loss function is designed to establish the goal of the agent's learning, and the model is trained by continuously reducing the gap between the target network and the agent network, so as to gradually identify and distinguish known mineralization, potential mineralization and non-mineralization in the interaction process of the agent and the environment, and delineate the metallogenic prospect area.

[0040] In step S30, a solution method for the Markov decision process is established. The Q-Learning method is used to directly learn the state-action value function Q(s, a), and the formula is as follows:

[0041] Q(s, a) = E[R t +γR t+1 +γ 2 R t+2 +…]

[0042] Wherein, E represents the mean; s represents the state; a represents the action; R represents the reward; and γ represents the discount factor, which represents the relationship between the current reward and the future reward.

[0043] Q-Learning adjusts the estimation of Q(s, a) through time difference, that is, Q(s, a) ← Q(s, a) + α[R + γmaxQ'(s, a) - Q(s, a)], wherein α is a control factor, and for the target Q' value, a neural network with the same architecture can be used to fit and represent. Thus, the agent can gradually distinguish known mineralization, potential mineralization and non-mineralization under the guidance of the reward signal, and the corresponding Q value gradually decreases.

[0044] The technical scheme provided by the application has the beneficial effects that: the application constructs a geometric graph structure and uses it as a state input of reinforcement learning to represent the complex spatial coupling relationship between the ore prospecting data and mineralization; the setting of the reinforcement learning reward mechanism guides the agent to comprehensively distinguish the mineralization potential based on multi-source mineralization information, which includes known mineralization information, potential mineralization information and prior mineralization knowledge, so that the prediction result is more in line with the mineralization regularity and the mineral prediction accuracy is improved; the agent constructed by the graph convolutional neural network can well utilize the powerful nonlinear fitting capability of deep learning to establish the nonlinear correlation between the geological ore prospecting data and the mineralization potential; the Markov chain provides a mathematical decision model for coupling the state-action-reward, and provides a solution for solving complex mineral prediction problems. The application can simultaneously consider multi-source geological information, integrate the mineralization information and mineralization knowledge into the Markov trajectory of the interaction between the agent and the environment, and in the training, optimize the difference between the Q value and the target Q' value of the Q-learning to enable the agent to gradually learn the benefits brought by the judgment of different states for mineralization, and the higher the benefits, the greater the possibility of mineralization, thereby improving the accuracy of mineral prediction.

[0045] The application will be described in detail below according to specific embodiments, but the protection scope of the application is not limited to the following embodiments.

[0046] Please refer to Figure 1 The specific embodiment provides a mineral prediction method based on graph reinforcement learning under geological constraints. The geometric graph provides a solution for modeling complex geological environments, the reward mechanism of reinforcement learning under geological constraints, the agent constructed by the graph convolutional neural network, and the interaction of the Markov chain, which makes deep interaction of geological knowledge and models possible. The model can simultaneously consider multi-angle geological mineralization information composed of known mineralization information, potential mineralization information and geological knowledge, thereby improving the accuracy of the mineral prediction model. The specific steps are as follows:

[0047] S1, constructing a geological ore prospecting data set according to a mineralization model, converting vector geological data into raster data using GIS technology, and simultaneously interpolating exploration geochemical data into a raster image. Based on the above process, the construction of the geological ore prospecting data set can be completed. Set a distance threshold to construct a geometric graph data structure as the state of graph reinforcement learning. Taking a certain research area as an example, set the buffer ring distance of the key ore-controlling elements to 1000 meters, simultaneously interpolate the 39-dimensional exploration geochemical data into a 1000m x 1000m raster image, and set a distance threshold of 3000 meters to construct the state in the graph reinforcement learning.

[0048] S2, constructing a reward mechanism of graph reinforcement learning and integrating geological knowledge (such as mineralization regularity) into it. For example Figure 2As shown, the reward mechanism of graph reinforcement learning can be modeled in three parts, which are known mineralization information, potential mineralization information and geological knowledge. Taking a certain research area as an example, the three types of reward information can be expressed as follows: (1) The mineralization position of known mineralization information as a direct index, the sample is divided into labeled mineralization set D1 and unlabeled set D0. If the state belongs to D1 (and is judged as known mineralization), then R1 = 1. If the state belongs to D0 (and is judged as unknown mineralization), then R1 = 0. Otherwise, R1 = -1. R1 aims to guide the model to preferentially learn the known mineralization information; (2) The potential mineralization information contained in the prospecting data set is mined by using the graph autoencoder. The reconstruction error is calculated by reconstructing the graph structure through low-dimensional embedding, which is used to evaluate the degree of geological anomaly. The node features X and the adjacency matrix A are encoded by using the graph convolutional network to obtain the low-dimensional embedding Z. The adjacency matrix A' is reconstructed by using Z, and the reconstruction error E = ‖A-A'‖ is calculated. E is normalized, and the normalized error value is R0, which is used to reflect the potential geological anomaly strength of the current state. The structure of the graph autoencoder model is as follows: firstly, the input features are mapped from the input data dimension to 64, 32 and 8 dimensions by three layers of graph convolution in turn, and normalization and ReLU activation are added after each layer to obtain the latent representation of the graph. Then, the latent representation is gradually decoded back to the original dimension by three layers of symmetric graph convolution, realizing the reconstruction of the node features and mining the spatial anomaly structure of the data; (3) The geological constraint reward R c : Based on the spatial distribution rule of ore deposits and ore-controlling elements, the relationship between the ore deposit density p and the distance d to the ore-controlling element is described by using a power function: p = Cd (m-2) , where C is a constant and m is a singular index. The function value is calculated and normalized to obtain R c . The closer to the ore-controlling element, the larger R c , which ensures that the model result conforms to the known spatial rule of mineralization.

[0049] S3, an intelligent agent of graph reinforcement learning is constructed, which is used to fit the nonlinear relationship between complex geological environment (state) and reward and action (mineralization potential), so as to train the intelligent agent to carry out mineral prediction. Due to the spatial heterogeneity of graph structure data, the graph convolutional neural network is used to mine the data. The working principle of the graph convolutional neural network is H (l) = σ9D -1 / 2 AD -1 / 2 H (l-1) W (l-1) +b (l)), where A denotes the adjacency matrix, D represents the degree matrix, H is the output feature matrix, and the superscripts l and l-1 represent the current layer and the previous layer, respectively. σ is the activation function, and W and b correspond to the weight and bias, respectively. Taking a study area as an example, one graph convolution layer is set, the dimension of the hidden layer is set to 64, and the σ is set to the ReLU activation function. The normalized node features output by the convolution layer are aggregated into a graph-level vector through global pooling, and then two fully connected layers are set, the hidden dimension of the first layer is 32, and the output dimension of the second layer is 2 (the size of the action space, to judge the ore-forming and non-ore-forming), so that the Q value of the action can be obtained.

[0050] S4, a Markov decision chain of graph reinforcement learning is constructed to provide a function for state transition. Taking a study area as an example, all states are divided into two categories: a known gold deposit data set (D1) and a data set without known mineralization information (D0). For the states in D1, they are randomly sampled to ensure that the agent learns various mineralization information. For the states in D0, a subset S is first selected according to the spatial autocorrelation Moran index, and the similarity of the data in the subset to the previous state is quantified. If the agent's action is judged to be an ore-forming anomaly, the next state tends to be in a region with similar attributes. If it is judged to be non-mineralized, a region with greater attribute difference is preferred. This state transition function can help the agent learn rich ore-forming anomaly characteristics and expand coverage of different geological features, thereby improving the generalization and recognition ability of the model.

[0051] S5, the loss function of the model is designed, and the time difference method of Q-Learning is used to maximize the reward obtained by the agent. That is, the graph convolutional neural network of the agent is used to fit the current Q value, and the target graph convolutional neural network is used to fit the Q value of the next state. The loss function L of the final graph reinforcement learning model training can be represented as:

[0052] L = E[(R t + γmaxQ'(s t+1 ,a t+1 )- Q(s t ,a t )) 2 ]

[0053] In the formula, t and t+1 represent the time sequence on the Markov chain; E represents the mean; R t represents the reward at time t; s t+1 represents the state at time t+1; a t+1 represents the action at time t+1; s t represents the state at time t; and a trepresents the action at time t; γ represents the discount factor, which balances the relationship between current reward and future reward; Q represents the Q value of the agent network fitting; Q′ represents the Q value of the target network fitting.

[0054] In addition, in the interaction between the agent and the environment in graph reinforcement learning, not only is the minimum value of the training loss function involved, but also the learning interaction strategy. This requires the model to freely explore trial and error in the early stage, without having to choose the Q-value action with the maximum value each time. This is called the ε-Greedy strategy. This strategy allows the agent to freely explore and randomly select actions in the early stage of the interaction to test the rewards it will receive, thereby continuously accumulating experience. As the interaction progresses, it will gradually standardize its own actions and gradually choose strategies that can obtain higher returns. Taking a certain research area as an example, in the first 90% of the interaction fragments in each Markov chain of the interaction, the agent's free exploration degree is linearly reduced from 1 to 0.1, and then remains unchanged, ensuring sufficient experience accumulation and training of the model's loss function. Finally, Q (known mineralization, a t =1)>Q(potential mineralization,a t =1)>Q(unmineralized,a t =1), the state data of each graph structure in the study area is input into the trained intelligent agent network, and the possibility of mineralization can be quantified according to the size of the Q value.

[0055] The graph reinforcement learning mineral prediction method based on geological constraints provided in this embodiment adopts a continuous trial and error and optimization strategy in the exploration data of the graph structure through the exploration-utilization mechanism of reinforcement learning, fully mining known mineralization information and capturing potential mineralization anomaly signals; at the same time, existing geological prior knowledge is incorporated into reward feedback as a constraint, guiding the model to prioritize areas that conform to mineralization laws, and promoting deep interaction between geological knowledge and artificial intelligence models. This embodiment realizes multidimensional mineralization decision-making under the deep coupling of known mineralization information, potential mineralization information and geological knowledge, which not only effectively narrows the scope of mineral exploration, but also increases the probability of discovering new mineral deposits, providing an efficient and reliable technical method for mineral prediction in complex geological environments.

[0056] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A mineral prediction method based on graph reinforcement learning based on geological constraints, characterized by: The steps include: S10, construct the state, reward and action in graph reinforcement learning. The construction method is: S11, constructing a geological prospecting dataset based on the metallogenic system and establishing the state of the graph reinforcement learning environment from it; S12, constructing reward feedback for graph reinforcement learning. In this construction process, geological constraints are added to enable the environment to reward the agent based on geological knowledge constraints. S13, constructing a graph reinforcement learning agent to intelligently interact with the environment. Based on the state and rewards of the environment feedback, it makes action decisions to determine whether there is mineralization potential, thereby fitting the nonlinear relationship between geological prospecting data and mineralization potential. S20, constructing a Markov chain decision process for graph reinforcement learning to achieve deep coupling between the agent and the environment; S30, designs a loss function to establish a goal for the agent's learning. The model is trained by continuously reducing the gap between the target network and the agent network. During the interaction between the agent and the environment, known mineralization, potential mineralization, and non-mineralization are gradually identified and distinguished, and prospective mineralization areas are delineated.

2. The method for mineral prediction based on graph reinforcement learning of geological constraints according to claim 1, characterized in that: In step S11, the geological prospecting dataset is rasterized, and on this basis, the raster data is converted into a geometric graph structure according to a set distance threshold as the state of graph reinforcement learning.

3. The method for mineral prediction based on graph reinforcement learning based on geological constraints according to claim 2, characterized in that: GIS technology is used to rasterize geological prospecting datasets.

4. The method for mineral prediction based on graph reinforcement learning based on geological constraints according to claim 1, characterized in that: In step S12, known mineralization information, potential mineralization information and geological knowledge are used as reward signals and deeply coupled, and together serve as the standard for the intelligent agent to judge whether the state has mineralization potential.

5. The method for mineral prediction based on graph reinforcement learning based on geological constraints according to claim 1 or 4, characterized in that: The geological knowledge includes mineralization laws.

6. The method for mineral prediction based on graph reinforcement learning of geological constraints according to claim 1, characterized in that: The agent adopts a graph convolutional neural network.

7. The method for mineral prediction based on graph reinforcement learning of geological constraints according to claim 1, characterized in that: In step S20, the interaction process between the intelligent agent and the environment is modeled as a Markov decision process, and a state transfer function is established, so that the intelligent agent can continuously make trial and error adjustments and ultimately seek the optimal mineralization potential prediction strategy.

8. The method for mineral prediction based on graph reinforcement learning of geological constraints according to claim 1, characterized in that: In step S30, the loss function L is: L=E[(R t +γmaxQ'(s t+1 ,a t+1 )-Q(s t ,a t )) 2 ] Where t and t+1 represent the time series on the Markov chain; E represents the mean; R t represents the reward at time t; s t+1 Indicates the state at time t+1; a t+1 represents the action at time t+1; s t Indicates the state at time t; a t represents the action at time t; γ represents the discount factor, which balances the relationship between current reward and future reward; Q represents the Q value of the agent network fitting; Q' represents the Q value of the target network fitting.

Citation Information

Cited By

  • Mineral development and ecological protection collaborative zoning method and device based on deep reinforcement learning, electronic equipment and storage medium

    CN121303904A