Geological guarantee intelligent planning mining system based on deep reinforcement learning
The intelligent planning and mining system for geological support based on deep reinforcement learning solves the problems of static fragility of geological support, short-sighted planning, and system fragmentation in mining, and realizes efficient and safe mining and system self-evolution throughout the entire life cycle of the mine.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SPACE VITE TECHNOLOGY DEVELOPMENT CO LTD
- Filing Date
- 2025-08-13
- Publication Date
- 2026-05-29
AI Technical Summary
Modern large-scale mining suffers from static and vulnerable geological conditions, short-sighted and non-strategic intelligent planning, and fragmented and open-loop mining systems, leading to a disconnect between planning and execution, and low system efficiency and safety.
A geological protection intelligent planning and mining system based on deep reinforcement learning is adopted, including a geological protection basic module, an intelligent planning engine module, and a closed-loop execution system module. A high-fidelity dynamic digital twin environment is constructed, and a long-term deep reinforcement learning decision-making agent for geological risk perception is developed to achieve adaptive and collaborative execution between planning and reality.
It has enhanced the dynamism and reliability of geological protection, achieved the unity of economic benefits and geological safety throughout the entire life cycle of the mine, broken down the information barriers between planning and execution, and realized the self-evolution of the system and efficient and safe mining.
Smart Images

Figure CN121031307B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent mining technology, and more specifically to an intelligent planning and mining system for geological support based on deep reinforcement learning. Background Technology
[0002] Modern large-scale mining is a complex systems engineering project characterized by high risk, high investment, and long cycle. Its efficiency and safety are highly dependent on the scientific nature of decision-making. However, existing technological systems face three fundamental dilemmas on the path to achieving true intelligent mining:
[0003] First, the static nature and vulnerability of geological assurance. Currently, geological assurance mainly relies on static three-dimensional geological models built based on sparse drilling data. These models are essentially discrete snapshots of geological bodies, containing unquantified uncertainties. Furthermore, they are slow to update, rely on manual intervention, and cannot dynamically reflect new information during the mining process. This causes geological assurance to become a static assumption, creating hidden dangers for planning.
[0004] Second, intelligent planning suffers from short-sightedness and lack of strategic thinking. Existing intelligent planning is mostly based on short-term local optimization using flawed geological models. Even when deep reinforcement learning is introduced, it still learns short-sighted strategies because it does not incorporate geological constraints. For example, prioritizing the mining of rich ore blocks can lead to stress concentration in the pillars or the overburdening of low-grade resources. It lacks consideration of geomechanical laws and is not truly intelligent.
[0005] Third, the fragmentation and open-loop nature of the mining system. An information gap exists between planning and implementation departments; on-site geological deviations require delayed manual adjustments; and the lack of an automatic, real-time closed-loop feedback mechanism leads to a disconnect between plans and reality, resulting in low system efficiency and safety.
[0006] Based on this, the present invention proposes a geological support intelligent planning and mining system based on deep reinforcement learning to solve the above problems. Summary of the Invention
[0007] To overcome the aforementioned deficiencies of the prior art, this invention provides a geological protection intelligent planning and mining system based on deep reinforcement learning to solve the problems existing in the background art.
[0008] This invention provides the following technical solution: a geological support intelligent planning and mining system based on deep reinforcement learning, and a geological support basic module for constructing a high-fidelity, dynamic digital twin environment oriented towards deep reinforcement learning;
[0009] The intelligent planning engine module is used to develop a long-term deep reinforcement learning decision-making intelligent agent for geological risk perception.
[0010] The closed-loop execution system module is used to achieve adaptive and collaborative execution of the bidirectional evolution of planning and reality.
[0011] Through the collaborative architecture of the three major modules, the core pain points of traditional mining systems, such as static fragility of geological support, short-sighted and non-strategic planning, and fragmented and open-loop systems, are fundamentally addressed. This enables dynamic cognition, far-reaching planning, and closed-loop self-evolution throughout the entire life cycle of the mine, providing a systematic solution for efficient and safe mining.
[0012] As a further aspect of the present invention: the geological support infrastructure module includes:
[0013] Based on the Geological Uncertainty Implicit Neural Field (GUINF) continuous probability representation unit, using parameters... Defined depth coordinate network Receive spatiotemporal coordinates Output the key geological attribute vector of this point. joint probability distribution ;
[0014] The gradient-flow-based dynamic evolution and data fusion unit realizes the dynamic evolution of the digital twin environment by calculating the negative log-likelihood loss of new data and updating the network parameters through gradient descent.
[0015] By overcoming the limitations of static geological models through the continuous probabilistic representation of GUINF and combining it with the gradient flow dynamic update mechanism, geological assurance is upgraded from static assumptions to dynamic, quantitative, real-time evolutionary models, effectively reducing the planning risks caused by geological uncertainties.
[0016] As a further aspect of the present invention: the GUINF adopts a sinusoidal representation network architecture, and the activation function is... This makes the network and its derivatives of any order continuous and analytically solvable with respect to the input coordinates; its output layer is a mixture density network, outputting the parameter set of a Gaussian mixture model. The probability density of geological attributes satisfies .
[0017] Sine representation networks ensure the continuity and differentiability of geological attribute characterization and support direct coupling with physical solvers; hybrid density networks accurately quantify the uncertainty of geological attributes, providing reliable probabilistic basis for planning decisions and improving the accuracy and credibility of geological assurance.
[0018] As a further aspect of the present invention: the intelligent planning engine module includes:
[0019] Geologically-aware Markov Decision Process (MDP) building blocks formalize planning problems into tuples. The state space Includes global geological protection characterization, geological risk maps of key areas, and resource endowment status, providing room for action. For a composite hierarchical decision instruction, the reward function R includes long-term value gain, geological risk penalty, and information uncertainty penalty;
[0020] Long-term programming algorithm unit based on causal world model, learns the world model. Predict state transitions and reward dynamics, and imagine and optimize long-term trajectories in the world model.
[0021] Geological perception MDP deeply integrates geological constraints into the decision-making framework, and the causal world model supports long-term trajectory planning, breaking through the short-sightedness of traditional planning, enabling decisions to take into account both short-term benefits and long-term geological safety, and maximizing the value of the entire life cycle of the mine.
[0022] As a further aspect of the present invention: the reward function is specifically in the form of:
[0023]
[0024] in, These are the weighting coefficients for economic value, geological risk, and information uncertainty, respectively. For the net present value increment, This is a penalty item for geological risks. This is a penalty item for information uncertainty.
[0025] Multi-objective reward functions balance economic value, geological risks, and information acquisition, guiding agents to proactively avoid risks and reduce cognitive uncertainty while pursuing benefits, avoiding sacrificing long-term safety for short-term gains, and improving the scientific and robust nature of planning.
[0026] As a further aspect of the present invention: the geological risk penalty item The calculation formula is: ,in, This represents the weighting coefficient for the risk item. For indicator functions, This is the stress value. For rock yield stress, The volume integral of the mining area. This represents the total energy of the microseismic event.
[0027] By quantifying the volume of rock mass stress exceeding limits and microseismic energy, geological risks (such as rock mass instability and ground pressure disasters) can be accurately measured, enabling intelligent agents to proactively avoid high-risk areas in the planning process and significantly improve the safety of the mining process.
[0028] As a further aspect of the present invention: the closed-loop execution system module includes:
[0029] The strategy distillation and secure deployment unit distills the decision-making knowledge of the trained DRL agent into a lightweight student neural network. In this context, real-time deployment is achieved through supervised learning;
[0030] The twin-reality bidirectional evolution closed-loop unit enables data flow updates from reality to the twin and proactive exploration from the twin to reality, triggering supplementary exploration tasks when key decisions depend on areas with high uncertainty.
[0031] Strategy distillation solves the problem of real-time deployment of complex models; bidirectional evolutionary closed loop breaks down the information barriers between planning and execution, enabling real-time data feedback and proactive inquiry, allowing the system to continuously learn and optimize from practice, and improve its adaptability to the actual mining environment.
[0032] As a further aspect of the present invention: a geologically sound intelligent planning and mining method based on deep reinforcement learning, comprising the following steps:
[0033] Step S1: Construct a high-fidelity, dynamic digital twin environment for deep reinforcement learning, specifically including continuous probability representation based on implicit neural fields of geological uncertainty and dynamic evolution and data fusion based on gradient flow;
[0034] Step S2: Develop a long-term deep reinforcement learning decision-making agent for geological risk perception, specifically including constructing a Markov decision process for geological perception and adopting a long-term planning algorithm based on a causal world model.
[0035] Step S3: Achieve adaptive and collaborative execution of the two-way evolution of planning and reality, specifically including strategy distillation and secure deployment, and establishing a closed loop for the two-way evolution of twin reality.
[0036] It provides a complete end-to-end technology path, covering everything from environment construction and intelligent agent development to closed-loop execution, ensuring that the system can be implemented and systematically solving the problems of poor dynamism, short-sighted planning, and insufficient coordination in traditional mining methods.
[0037] As a further aspect of the present invention: step S1 includes:
[0038] S1.1: The mine is implicitly represented by the GUINF network, and the joint probability distribution of geological attributes at any spatiotemporal coordinate is output.
[0039] S1.2: The sinusoidal representation of the network architecture enables the network to have continuous differentiability.
[0040] S1.3: Output Gaussian mixture model parameters through a mixture density network;
[0041] S1.4: Negative log-likelihood loss based on new observation data Through gradient descent Update network parameters.
[0042] Step S1 ensures that the digital twin environment has high fidelity and dynamic evolution characteristics, can absorb new data in real time and quantify uncertainty, and provide an accurate and reliable geological understanding basis for subsequent planning and decision-making.
[0043] As a further aspect of the present invention: step S2 includes:
[0044] S2.1: Define the MDP state space as including global geological characterization, local risk map and resource status, the action space as a hierarchical discrete structure, and the reward function as a multi-objective weight;
[0045] S2.2: Train a Transformer-based world model, and optimize long-cycle trajectories in the world model using the cross-entropy method, with the planning vision set to 240 monthly decision steps.
[0046] Step S2 enables the agent to deeply perceive geological constraints and conduct long-term planning, avoiding short-term optimization traps and ensuring that planning decisions achieve the optimal balance between efficiency and safety throughout the entire life cycle of the mine.
[0047] The technical effects and advantages of this invention are as follows:
[0048] (1) Fundamentally improve the level of geological security: Through differentiable probability digital twins, this invention elevates geological security from a static and qualitative concept to a dynamic, quantitative, and reliable commitment that can evolve in real time with mining activities, providing an unprecedented solid foundation for all planning decisions.
[0049] (2) Achieving intelligent planning: Through the DRL framework for geological risk perception, this invention enables planning decisions to transcend short-term profit orientation, and to proactively manage and avoid long-term geological risks, achieving a deep unity of economic benefits and geological safety, and making it possible to maximize the value of the entire life cycle of the mine.
[0050] (3) A closed-loop self-evolving mining system was constructed: Through the bidirectional evolution and active exploration mechanism of twin reality, this invention breaks down the barriers between planning and execution, and creates an intelligent system that can learn from practice, actively improve its own cognition, and achieve the co-evolution of planning and reality. Attached Figure Description
[0051] The invention will now be further described with reference to the accompanying drawings.
[0052] Figure 1 A schematic diagram of the overall architecture of the intelligent planning and mining system for geological protection based on deep reinforcement learning provided in an embodiment of the present invention;
[0053] Figure 2 This is a flowchart illustrating the bidirectional evolutionary closed loop of twin reality in this embodiment of the invention.
[0054] Figure 3 This is a schematic diagram of the framework of the geological perception Markov decision-making process in an embodiment of the present invention:
[0055] Figure 4 This is a system block diagram of the closed-loop execution system module in an embodiment of the present invention. Detailed Implementation
[0056] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention. Example
[0057] like Figure 1 As shown, this embodiment provides a geological support intelligent planning and mining system based on deep reinforcement learning. The system includes three core modules: a geological support basic module, an intelligent planning engine module, and a closed-loop execution system module.
[0058] The Geological Support Foundation Module is responsible for building a high-fidelity, dynamic digital twin environment for deep reinforcement learning. This module includes two key units:
[0059] Continuous probability representation unit based on Geological Uncertainty Implicit Neural Field (GUINF):
[0060] like Figure 2 As shown, this unit uses a single parameter... Defined depth coordinate network To implicitly represent the entire mine. The network receives arbitrary spatiotemporal coordinates. As input, output the vector of all key geological attributes for that point. joint probability distribution .vector It includes two types of key information: resource attributes (such as ore grade and density) and engineering attributes (such as rock quality indicators RMR, elastic modulus, Poisson's ratio, cohesion C, and internal friction angle φ).
[0061] The GUINF network uses a sinusoidal representation network (SIREN) architecture, and its activation function is: This makes the network Its derivatives of any order are continuous and analytically solvable with respect to the input coordinates. This property allows for direct coupling of the digital twin with a physical solver based on partial differential equations.
[0062] To provide true geological assurance, namely a reliable quantification of uncertainty, the output layer of GUINF is a mixture density network (MDN), whose output is the parameter set defining a Gaussian mixture model (GMM). Therefore, at any point Any geological attribute The probability density is:
[0063] Gradient-flow-based dynamic evolution and data fusion unit:
[0064] This unit makes the digital twin a "living" entity, capable of dynamically absorbing new data generated during the mining process. When on-site mining reveals new geological information, these new observational data, with precise spatiotemporal tags, become readily available. This will be used for online fine-tuning of the GUINF network. The negative log-likelihood loss will be calculated using the new data in the current model:
[0065]
[0066] And update the network parameters using gradient descent:
[0067] The intelligent planning engine module is responsible for developing a long-term deep reinforcement learning decision-making agent for geological risk perception. For example... Figure 3 As shown, this module includes:
[0068] Geologically-aware Markov Decision Process (MDP) Building Blocks:
[0069] The intelligent planning problem is formalized as a complex geological sensing MDP, whose tuples are defined as follows: State space It is a high-dimensional, structured vector that comprehensively describes the mine at any given time. The situation includes: global geological safeguard representation: digital twin network parameters The features of the entity itself, or its feature embedding on a predefined grid; Geological risk map of key areas: spatial distribution of key physical fields obtained by querying through digital twins and calculating through stress field solvers; Resource endowment status: spatial distribution of remaining recoverable reserves, grade distribution and its uncertainty.
[0070] reward function Designed as a multi-objective, risk-punishment model:
[0071]
[0072] Among them, geological risk penalty items The calculation formula is:
[0073]
[0074] Long-term programming algorithm unit based on causal world model:
[0075] The agent first learns a world model The model is able to predict the approximate dynamics of state transitions and reward functions. The agent then performs massively parallel long-period trajectory visualization within its world model, using gradient-based trajectory optimization algorithms or Monte Carlo tree search to select the action sequence with the optimal long-term cumulative reward.
[0076] The closed-loop execution system module is responsible for achieving adaptive and collaborative execution of the bidirectional evolution of planning and reality. For example... Figure 4 As shown, this module includes:
[0077] Strategy Distillation and Secure Deployment Unit: Employing strategy distillation technology, the wisdom of the teacher model is "distilled" into a lightweight student neural network. In this approach, the teacher model makes decisions on a massive number of states within a digital twin, forming a huge dataset of (states, optimal actions), which is then used to train the student model using supervised learning.
[0078] Twin Reality Two-Way Evolution Closed-Loop Unit: Establishing a two-way information flow: Data flow from reality to the twin: All new geological data from actual mining activities will be used to update the digital twin in real time and smoothly; Exploration flow from the twin to reality: When the DRL agent discovers that the quality of a key decision is highly dependent on a high-uncertainty area in the digital twin, it will actively trigger the "active exploration" task to calculate the information value (VoI) of supplementary exploration in the uncertain area. Example
[0079] This embodiment provides a geological protection intelligent planning and mining method based on deep reinforcement learning. The specific implementation steps are as follows:
[0080] Step 1: Construct a high-fidelity, dynamic digital twin environment for deep reinforcement learning.
[0081] First, initial geological data for model training was collected and processed. Data sources included: historical borehole data: core logging data from all boreholes during the exploration and production exploration phases of the mine, totaling 2,500 boreholes, yielding spatial coordinates of approximately 150,000 samples. And corresponding ore grade, lithology, RMR value, and other attributes. Geological profile maps: Digitize the existing 25 key geological profile maps, extract the spatial location information of major lithological contact zones, faults, and alteration zones, and generate approximately 30,000 control point data. Existing stope exposure data: Organize the geological sketches, ore grade acceptance records, and ground pressure monitoring data of the mined areas over the past 5 years, align them with the 3D model of the stopes, and generate approximately 50,000 data points with accurate spatiotemporal labels. And data points with corresponding attributes. Geophysical data: If available, 3D seismic exploration or resistivity tomography data can be introduced as soft constraints to provide regional geological structure trend information.
[0082] Secondly, the collected multi-source heterogeneous data underwent preprocessing: Coordinate System 1: All data were transformed to a unified mining coordinate system. Data Cleaning: Outlier data points with incomplete records or obvious errors were removed. Normalization: All geological attribute values (such as grade, RMR, elastic modulus, etc.) were subjected to min-max normalization and scaled to [the appropriate value]. Within a given interval, this facilitates stable training of the neural network. Constructing the training set: Integrating the processed data points into a training set with a unified format. ,in , For 4-dimensional spatiotemporal coordinates (for static data, time...) Set to 0). This is the normalized geological attribute vector for that point.
[0083] Then, the GUINF network is initialized and trained. Network architecture: The network uses an 8-layer fully connected structure, with 256 neurons in each layer, and the activation function is... The input layer receives position-encoded 4D coordinates. The output layer is a hybrid density network, outputting for each geological attribute. Parameters of each Gaussian component (mixing coefficient) mean ,Difference Pre-training process: using the training set The network is pre-trained. The loss function is the sum of the negative log-likelihoods of all training data points:
[0084]
[0085] Formula explanation: The entire training set The total loss function on. :From the training set arrive Sum all the data points. Model for the first Data points The log-likelihood. The goal of training is to find a set of parameters. This maximizes the total probability (i.e., total likelihood) of all known data points, which is equivalent to minimizing the negative log-likelihood loss.
[0086] The Adam optimizer is used, with an initial learning rate of The batch size is 64, and the training is conducted for 10,000 epochs until the loss function converges, thus obtaining the initial digital twin environment.
[0087] Step 2: Develop a long-term deep reinforcement learning decision-making agent for geological risk perception.
[0088] First, define the specific parameters of the geological sensing MDP. State space. The dimension is set to 1024, and it is composed of three parts: Global geological representation (512 dimensions): The current GUINF network is placed in a fixed dimension. The system queries the grid to obtain the mean and variance of geological attributes for each grid point, then compresses them into a 512-dimensional feature vector using a small convolutional autoencoder. Local risk map (256-dimensional): For several key areas requiring decision-making, software such as FLAC3D is used to perform rapid stress simulation based on rock mechanics parameters obtained from GUINF queries, extracting key indicators such as maximum principal stress and stress concentration factor, and encoding them into a 256-dimensional vector. Resource status (256-dimensional): Information such as the spatial distribution, average grade, and total reserves of remaining oreable blocks is statistically analyzed and encoded into a 256-dimensional vector. Action space. The design employs a layered discrete structure: the upper layer involves the selection of 10 candidate stopes (discrete action 110); the middle layer involves the selection of three mining methods (such as stratified caving, segmented open-face mining, and backfilling) (discrete action 13); and the lower layer represents the monthly mining intensity (discrete action, categorized into low, medium, and high levels). Reward function weights are set based on the mine's long-term strategic goals. , , In geological risk penalties, (Penalty for volume instability) (Penalty for microseismic energy). Discount factor This is to encourage long-term planning.
[0089] Then, the causal world model is trained. The agent interacts with the digital twin environment constructed in step 1 to collect data. The model uses empirical data (state, action, next state, reward) to train a Transformer-based sequence model as a world model. This model has a hidden dimension of 512 and includes 8 attention heads and a 6-layer encoder-decoder structure.
[0090] Finally, long-term planning training is performed. The cross-entropy method (CEM) is used for trajectory optimization within the learned world model. The planning horizon is set to 240 monthly decision steps. At each decision point, the agent imagines and samples 1000 action sequences for the next 240 steps in parallel, evaluates the cumulative reward of each sequence, selects the top 10% (i.e., 100) of the best-performing sequences, and updates the sampling distribution for the next round using their mean and variance. After 5 iterations, the first action of the top-ranked action sequence in the updated distribution is executed.
[0091] Step 3: Achieve adaptive and collaborative execution of the two-way evolution of planning and reality.
[0092] First, policy distillation is performed. A pre-trained, computationally expensive CEM-based teacher model is used to infer from 1000 different states randomly sampled from a digital twin environment, collecting (state, optimal action) data pairs. These data pairs are then used to train a lightweight student network in a supervised learning manner. This student network is a 3-layer fully connected neural network (128 neurons per layer, ReLU activation), using cross-entropy loss for behavior cloning.
[0093] Then, a closed-loop system is deployed. The distilled student strategy network is deployed to the mine's central scheduling system, automatically generating and issuing the monthly mining plan at the beginning of each month. At the same time, an automated data acquisition channel is established to collect new data generated by underground production, measurement, geological survey, and monitoring systems (such as actual ore grade, lithological changes revealed in roadways, microseismic monitoring events, etc.) into the database in real time.
[0094] Finally, proactive exploration is implemented. A monitoring module continuously runs in the system background. This module periodically (e.g., weekly) assesses the geological uncertainties of key areas to be mined. It uses the Monte Carlo method to sample multiple sets (e.g., 100 sets) of possible geological realizations from the probability distribution of GUINF, and runs a planning agent for each realization to evaluate the value of the decision. When the system detects that the geological uncertainty in a certain area causes the expected NPV fluctuation range of subsequent planning to exceed a preset threshold, it automatically generates a "proactive exploration" task, sending supplementary exploration suggestions to the geological survey department. The suggestions include optimized borehole locations, depths, and target sampling densities.
[0095] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A geological support intelligent planning and mining system based on deep reinforcement learning, characterized in that: include: The geological support module is used to build a high-fidelity, dynamic digital twin environment for deep reinforcement learning. The intelligent planning engine module is used to develop a long-term deep reinforcement learning decision-making intelligent agent for geological risk perception. The closed-loop execution system module is used to achieve adaptive and collaborative execution of the two-way evolution of planning and reality; The geological support infrastructure module includes: Based on the Geological Uncertainty Implicit Neural Field (GUINF) continuous probability representation unit, using parameters... Defined depth coordinate network Receive spatiotemporal coordinates Output the key geological attribute vector of this point. joint probability distribution ; The gradient-flow-based dynamic evolution and data fusion unit realizes the dynamic evolution of the digital twin environment by calculating the negative log-likelihood loss of new data and updating the network parameters through gradient descent.
2. The intelligent planning and mining system for geological support based on deep reinforcement learning according to claim 1, characterized in that: The GUINF uses a sinusoidal representation network architecture, and the activation function is... Where u is the input variable of the hidden layer of the network, making the network and its derivatives of any order continuous and analytically solvable with respect to the input coordinates; its output layer is a mixture density network, outputting the parameter set of a Gaussian mixture model. The probability density of geological attributes satisfies Where V is the key geological attribute vector, C is the input spatiotemporal coordinates, and K is the total number of distribution components in the Gaussian mixture model. The first The weights, mean, and covariance of each component.
3. The intelligent planning and mining system for geological support based on deep reinforcement learning according to claim 1, characterized in that: The intelligent planning engine module includes: Geologically-aware Markov Decision Process (MDP) building blocks formalize planning problems into tuples. The state space Includes global geological protection characterization, geological risk maps of key areas, and resource endowment status, providing room for action. For the composite hierarchical decision instruction, P is the state transition probability, which represents the probability of transitioning to the next state after performing action A in state S. The reward function R includes long-term value gain, geological risk penalty and information uncertainty penalty. γ is the discount factor for future rewards, which is used to measure the current value of future rewards. Long-term programming algorithm unit based on causal world model, learns the world model. Predict state transitions and reward dynamics, and imagine and optimize long-term trajectories in the world model.
4. The intelligent planning and mining system for geological protection based on deep reinforcement learning according to claim 3, characterized in that: The reward function takes the following specific form: in, These are the weighting coefficients for economic value, geological risk, and information uncertainty, respectively. For the net present value increment, This is a penalty item for geological risks. This is a penalty item for information uncertainty. The state at the current time step t, This is the state at the next time step t+1.
5. The intelligent planning and mining system for geological protection based on deep reinforcement learning according to claim 4, characterized in that: Geological risk penalty items The calculation formula is: ,in, This represents the weighting coefficient for the risk item. For indicator functions, This is the stress value. For rock yield stress, The volume integral of the mining area. This represents the total energy of the microseismic event.
6. The intelligent planning and mining system for geological support based on deep reinforcement learning according to claim 1, characterized in that: The closed-loop execution system module includes: The strategy distillation and secure deployment unit distills the decision-making knowledge of the trained DRL agent into a lightweight student neural network. In this context, real-time deployment is achieved through supervised learning; The twin-reality bidirectional evolution closed-loop unit enables data flow updates from reality to the twin and proactive exploration from the twin to reality, triggering supplementary exploration tasks when key decisions depend on areas with high uncertainty.
7. A geological protection intelligent planning and mining method based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a high-fidelity, dynamic digital twin environment for deep reinforcement learning, specifically including continuous probability representation based on implicit neural fields of geological uncertainty and dynamic evolution and data fusion based on gradient flow; Step S2: Develop a long-term deep reinforcement learning decision-making agent for geological risk perception, specifically including constructing a Markov decision process for geological perception and adopting a long-term planning algorithm based on a causal world model. Step S3: Achieve adaptive and collaborative execution of the bidirectional evolution of planning and reality, specifically including strategy distillation and secure deployment, and establishing a closed loop for the bidirectional evolution of twin reality; Step S1 includes: S1.1: The mine is implicitly represented by the GUINF network, and the joint probability distribution of geological attributes at any spatiotemporal coordinate is output. S1.2: The sinusoidal representation of the network architecture enables the network to have continuous differentiability. S1.3: Output Gaussian mixture model parameters through a mixture density network; S1.4: Negative log-likelihood loss based on new observation data ,in For the geological attribute values of the new observation data, For the spatiotemporal coordinates of the new observation data, gradient descent is used. Update network parameters, including This is the learning rate.
8. The intelligent planning and mining method for geological protection based on deep reinforcement learning according to claim 7, characterized in that: Step S2 includes: S2.1: Define the MDP state space as including global geological characterization, local risk map and resource status, the action space as a hierarchical discrete structure, and the reward function as a multi-objective weight; S2.2: Train a Transformer-based world model, and optimize long-cycle trajectories in the world model using the cross-entropy method, with the planning vision set to 240 monthly decision steps.