An open-pit mine digital safety production management method, device and medium

By constructing digital twins and meta-reinforcement learning agents, blasting parameters were optimized, solving the problem of dynamic adjustment of blasting schemes in open-pit mines under complex geological conditions. This achieved synergistic optimization of safety and efficiency, and improved the adaptability and reliability of blasting schemes.

CN121389826BActive Publication Date: 2026-04-21CHANGCHUN GOLD DESIGN INST
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGCHUN GOLD DESIGN INST
Filing Date
2025-12-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing open-pit mine safety management methods cannot dynamically adjust blasting parameters under complex geological conditions, making it difficult to coordinate and optimize blasting schemes between safety thresholds and blasting efficiency, and failing to effectively solve the dynamic balance problem between vibration control and crushing effect.

Method used

By constructing a digital twin, a meta-reinforcement learning agent is used to explore different combinations of blasting parameters in the blasting parameter space. The physical and mechanical parameters of the rock mass are optimized by combining physical information neural networks to generate a blasting scheme that balances safety and efficiency. Through multiple rounds of simulation and parameter iterative optimization, an adaptive and reliable blasting scheme is generated.

Benefits of technology

It enables the optimization of blasting schemes under dynamic geological conditions, improves the adaptability and reliability of blasting schemes, ensures that the economy and safety of blasting effects are in a coordinated state, and provides accurate basis for safe production decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121389826B_ABST
    Figure CN121389826B_ABST
Patent Text Reader

Abstract

This invention discloses a digital safety production management method, equipment, and medium for open-pit mines, relating to the field of intelligent mining technology. The method includes: collecting rock mass physical and mechanical parameters and three-dimensional geometric data of the blasting face in the gold mining area, and constructing a digital twin as physical constraints; inputting a pre-set target blasting scheme into a meta-reinforcement learning agent, performing multiple rounds of simulation and deduction through a blasting scheme evaluation strategy, and generating simulation results; when the peak particle velocity in the simulation results exceeds a preset peak particle velocity threshold, generating parameter adjustment suggestions; feeding the parameter adjustment suggestions back to the meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed the preset peak particle velocity threshold, and then generating a blasting scheme. This invention, through the autonomous exploration of the meta-reinforcement learning agent in the blasting parameter space, can efficiently generate blasting scheme evaluation strategies that are both adaptive and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent mining technology, and in particular to a digital safety production management method, equipment and medium for open-pit mines. Background Technology

[0002] With the deepening application of digital twin technology in the mining industry, current open-pit mine safety production management can construct high-precision virtual models by integrating three-dimensional geological modeling with rock mechanics parameters. Existing technologies can achieve optimized simulation of blasting parameters and preliminary prediction of vibration propagation, and use finite element analysis methods to verify the safety of blasting schemes, providing technical support for mine digitalization.

[0003] However, existing methods have shortcomings when dealing with multi-objective optimization under complex geological conditions. The main problem is that they rely on static models and preset rules, and cannot dynamically adjust the blasting parameter optimization strategy based on real-time geological data. This results in a lack of adaptability to changes in rock mass conditions, difficulty in achieving synergistic optimization of safety threshold and blasting efficiency, and an inability to effectively solve the dynamic balance problem between vibration control and fracturing effect. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a digital safety production management method for open-pit mines to solve the problem that the safety and economy of blasting schemes are difficult to optimize in a coordinated manner due to the prediction bias of static models and the insufficient multi-objective dynamic optimization capabilities.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a digital safety production management method for open-pit mines, comprising: collecting rock mass physical and mechanical parameters and three-dimensional geometric morphology data of blasting faces in gold mining areas, and constructing a digital twin as physical constraints; setting a blasting parameter space in the digital twin, and initializing a meta-reinforcement learning agent in the digital twin; exploring the safety and efficiency results corresponding to different combinations of blasting parameters in the blasting parameter space through the meta-reinforcement learning agent, and obtaining a blasting scheme evaluation strategy; calling a preset target blasting scheme and inputting it into the meta-reinforcement learning agent, and performing multiple rounds of simulation and deduction through the blasting scheme evaluation strategy to generate simulation and deduction results; when the peak particle velocity in the simulation and deduction results exceeds a preset peak particle velocity threshold, generating parameter adjustment suggestions; feeding back the parameter adjustment suggestions to the meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed the preset peak particle velocity threshold, and generating a blasting scheme.

[0008] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the specific steps for collecting the physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting face in the gold mine mining area, and constructing a digital twin using these as physical constraints, are as follows.

[0009] The physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting face are used as physical constraints and input into the physical information neural network.

[0010] The encoder in the physical information neural network converts the three-dimensional geometric data of the blasting face into a latent space feature vector.

[0011] The latent space feature vector and rock mass physical and mechanical parameters are fused and calculated using the decoder in the physical information neural network to generate predicted stress distribution and displacement field data.

[0012] The predicted stress distribution and displacement field data are compared with the real-time monitored slope displacement and vibration data to obtain residual values. The physical constraint weight coefficients of the physical information neural network are then optimized based on the residual values.

[0013] The optimized physical constraint weight coefficients are solidified into a digital twin.

[0014] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the blasting parameter space includes borehole spacing, borehole depth, charge quantity, and detonation delay sequence.

[0015] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the specific steps for setting a blasting parameter space in the digital twin and initializing the meta-reinforcement learning agent in the digital twin are as follows:

[0016] Based on the predicted stress distribution and displacement field data, the joint probability distribution function of borehole spacing, borehole depth, charge amount and detonation delay sequence is calculated;

[0017] Based on the joint probability distribution function, the marginal probability distributions of borehole spacing, borehole depth, charge amount and detonation delay sequence parameters are extracted respectively. The cumulative distribution function is calculated for the marginal probability distributions of each parameter to obtain the high probability density interval as the search range of the corresponding parameter, thus completing the setting of the blasting parameter space.

[0018] After the setup is complete, the high-level meta-policy network and the low-level execution policy network are embedded into the meta-reinforcement learning agent and initialized.

[0019] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the step of exploring the safety and efficiency results corresponding to different combinations of blasting parameters in the blasting parameter space through a meta-reinforcement learning agent to obtain a blasting scheme evaluation strategy involves the following specific steps:

[0020] By using a high-level meta-policy network within a meta-reinforcement learning intelligent body, feature extraction and target reasoning are performed on the predicted stress distribution and displacement field data to generate exploratory sub-targets.

[0021] Based on the safety constraints in the exploration sub-objectives, the underlying execution policy network in the meta-reinforcement learning intelligent body is used to sample the safety constraints in the blasting parameter space to generate a candidate blasting parameter combination set.

[0022] By using a meta-reinforcement learning agent to predict the safety and efficiency results of candidate blasting parameter combinations, the predicted safety and efficiency indicators are obtained.

[0023] The reward value of each candidate blasting parameter combination in the candidate blasting parameter combination set is calculated based on the predicted safety and efficiency indicators, and the candidate blasting parameter combination with the largest reward value is taken as the current blasting parameter combination.

[0024] The current combination of blasting parameters is input into the digital twin to update and optimize the parameters of the high-level meta-strategy network until the high-level meta-strategy network parameters reach the convergence condition. Then, the high-level meta-strategy network parameters are solidified into the blasting scheme evaluation strategy.

[0025] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the steps of calling a preset target blasting scheme input meta-reinforcement learning agent, performing multiple rounds of simulation deduction through a blasting scheme evaluation strategy, and generating simulation deduction results are as follows:

[0026] The preset target blasting scheme is input into the meta-reinforcement learning agent, which analyzes and extracts features from the preset target blasting scheme to generate a scheme feature vector.

[0027] By using a blasting scheme evaluation strategy, the scheme feature vectors are simulated and deduced in multiple rounds in parallel to obtain a simulation and deduction instruction set;

[0028] Parallel simulation and deduction of the simulation and deduction instruction set are performed using a digital twin, and the simulation indicators for each round are output.

[0029] The performance stability and robustness of each round of simulation indicators are evaluated and quantified by a multi-index fusion analysis algorithm, generating simulation results.

[0030] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the step of generating parameter adjustment suggestions when the peak particle velocity in the simulation results exceeds a preset peak particle velocity threshold is as follows:

[0031] Extract the peak particle velocity from the simulation results. When the peak particle velocity exceeds the preset peak particle velocity threshold, trigger the parameter adjustment signal.

[0032] Based on the parameter adjustment signal, the correlation characteristic analysis of the current blasting parameter combination is initiated, and the analysis results are matched with the historical exploration data of the meta-reinforcement learning agent to generate matching results;

[0033] The matching results are evaluated and optimized through a comprehensive analysis mechanism, generating parameter adjustment suggestions.

[0034] As a preferred embodiment of the digital safety production management method for open-pit mines described in this invention, the step of feeding back parameter adjustment suggestions to a meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed a preset peak particle velocity threshold, generating a blasting plan, and inputting the blasting plan to the gold mine production scheduling terminal, is as follows:

[0035] Input the parameter adjustment suggestions into the meta-reinforcement learning agent to update the high-level meta-policy network parameters in the meta-reinforcement learning agent, and obtain the updated high-level meta-policy network parameters.

[0036] The current blasting parameter combination is iteratively optimized based on the updated high-level meta-strategy network parameters until the peak particle velocity does not exceed the preset peak particle velocity threshold. Then, the optimized blasting parameter combination is generated and input as the blasting scheme to the gold mine production scheduling terminal.

[0037] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the digital safety production management method for open-pit mines as described in the first aspect of the present invention.

[0038] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the digital safety production management method for open-pit mines as described in the first aspect of the present invention.

[0039] The beneficial effects of this invention are as follows: Through the autonomous exploration of the blasting parameter space by the meta-reinforcement learning agent, a blasting scheme evaluation strategy with both adaptability and reliability can be efficiently generated; through the interactive learning between the agent and the digital twin, a comprehensive performance evaluation of different combinations of blasting parameters under dynamic geological conditions can be achieved, generating decision rules that combine dynamic control and crushing effect, which can quickly identify the parameter balance point, improve the adaptability and reliability of the blasting scheme, provide accurate decision-making basis for mine safety production, and at the same time ensure that the economy and safety of the blasting effect are in a coordinated state. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 A flowchart for digital safety production management methods in open-pit mines.

[0042] Figure 2 A flowchart for building a digital twin.

[0043] Figure 3 A flowchart for obtaining an evaluation strategy for blasting schemes.

[0044] Figure 4 This is a flowchart for simulation deduction and parameter iterative optimization. Detailed Implementation

[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0046] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0047] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0048] Reference Figures 1-4As an embodiment of the present invention, this embodiment provides a digital safety production management method for open-pit mines, comprising the following steps:

[0049] S1. Collect the physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting operation face in the gold mining area, and use them as physical constraints to construct a digital twin.

[0050] S1.1: Input the physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting face as physical constraints into the physical information neural network;

[0051] It should be noted that the physical and mechanical parameters of the rock mass refer to the quantitative indicators of rock mass density, compressive strength and elastic modulus obtained through geological drilling exploration; the three-dimensional geometric morphology data of the blasting operation face refers to the spatial topological data of the borehole distribution coordinates and free surface contour information obtained through a three-dimensional laser scanner.

[0052] It should be noted that the pre-training process of the physical information neural network is as follows: physical constraints are constructed using historical rock mass physical and mechanical parameters and historical blasting operation face three-dimensional geometric morphology data. The physical constraints guide the optimization of the weights of the physical information neural network, enabling the physical information neural network to learn the rock mass blasting response law, forming a digital twin that can simulate the blasting process, and obtaining the pre-trained physical information neural network.

[0053] Specifically, the rock mass density, compressive strength, and elastic modulus from the rock mass physical and mechanical parameters, together with the borehole distribution coordinates and free surface contour information from the three-dimensional geometric morphology data of the blasting operation face, constitute a set of physical constraints. This set of physical constraints is applied to the training process of the physical information neural network. Through the constraint loss function, the rock mass physical and mechanical parameters and the three-dimensional geometric morphology data of the blasting operation face are embedded into the weight update process of the physical information neural network, enabling the physical information neural network to simulate the blasting response behavior of rock mass.

[0054] S1.2: The encoder in the physical information neural network converts the three-dimensional geometric data of the blasting face into a latent space feature vector;

[0055] Specifically, the encoder includes convolutional layers, pooling layers, and fully connected layers. The convolutional layers extract local features from the borehole distribution coordinates and free surface contour information using a sliding window, capturing local spatial features. The pooling layers compress the dimensions of the local spatial features while preserving the topological relationships, generating a dimensionality-reduced feature map. The fully connected layers linearly combine all nodes of the dimensionality-reduced feature map with the weight matrix, and then perform a nonlinear transformation using an activation function to generate a latent space feature vector.

[0056] It should be noted that the weight matrix was optimized using the backpropagation algorithm during the initial training process of the physical information neural network.

[0057] S1.3: The latent space feature vector and rock mass physical and mechanical parameters are fused and calculated using the decoder in the physical information neural network to generate predicted stress distribution and displacement field data. The expression is as follows:

[0058] ;

[0059] ;

[0060] ;

[0061] in, Represents spatial coordinates The predicted stress value at the location; Represents a spatial coordinate vector; This refers to a mapping function or subnetwork in a neural network decoder specifically used to generate stress distribution; This represents the latent space feature vector extracted from the encoder; This represents a vector of physical and mechanical parameters of the rock mass, including rock mass density, compressive strength, and elastic modulus. This represents the output layer weight matrix; Indicates the activation function; Represents the hidden layer weight matrix; This represents a vector concatenation operation; Represents the hidden layer bias vector; Indicates the output layer bias term; Represents spatial coordinates The predicted displacement value at the location; This represents the weight matrix of the output layer; This represents the weight matrix of the hidden layer; Represents the predicted latent space feature vector; Represents the vector of predicted rock mass physical and mechanical parameters; This represents the bias vector of the hidden layer; This represents the bias scalar of the output layer.

[0062] It should be noted that the weight matrix refers to the two-dimensional parameter structure used in the fully connected layer of a high-level meta-policy network to store the connection weights between neurons, and is used for linear transformation operations during forward propagation.

[0063] S1.4: Compare the predicted stress distribution and displacement field data with the real-time monitored slope displacement and vibration data to obtain the residual value, and optimize the physical constraint weight coefficient of the physical information neural network based on the residual value.

[0064] Specifically, the predicted stress distribution and displacement field data generated by the physical information neural network decoder are spatially and temporally aligned with the slope displacement and vibration data monitored in real time by the GNSS receiver and vibration sensor network to generate aligned predicted data and aligned monitoring data. The residual value is obtained by comparing the difference between the corresponding values ​​of the aligned predicted data and the aligned monitoring data. Based on the residual value, the physical constraint weight coefficients in the physical information neural network are optimized and adjusted to enable the physical information neural network to more accurately simulate the rock mass mechanical behavior.

[0065] It should be noted that the real-time monitored slope displacement and vibration data refers to the real-time synchronous acquisition of millimeter-level slope displacement data and blasting vibration waveform data through a GNSS receiver and vibration sensor network.

[0066] S1.5: Solidify the optimized physical constraint weight coefficients into a digital twin.

[0067] Specifically, the optimized physical constraint weight coefficients are applied to the parameter configuration of the physical information neural network, enabling the digital twin to acquire stable physical constraint characteristics, thereby completing the solidification of the digital twin.

[0068] S2. Set up a blasting parameter space in the digital twin and initialize the meta-reinforcement learning agent in the digital twin. Explore the safety and efficiency results corresponding to different combinations of blasting parameters in the blasting parameter space through the meta-reinforcement learning agent to obtain the blasting scheme evaluation strategy.

[0069] S2.1: The blasting parameter space includes borehole spacing, borehole depth, charge amount, and detonation delay sequence.

[0070] It should be noted that the borehole spacing refers to the horizontal distance between the center points of two adjacent boreholes on the blasting operation surface, the borehole depth refers to the vertical depth of the borehole from the opening to the bottom, the charge amount refers to the total weight of explosives loaded in a single borehole, and the detonation delay sequence refers to the order of detonation time delays set for different boreholes or borehole groups in multi-hole blasting.

[0071] S2.2: Based on the predicted stress distribution and displacement field data, calculate the joint probability distribution function of borehole spacing, borehole depth, charge amount, and detonation delay sequence. The expression is:

[0072] ;

[0073] in, Denotes the joint probability density function; Indicates the spacing between boreholes; Indicates the drilling depth; Indicates the amount of explosives; Indicates the detonation delay sequence; This represents the predicted displacement field data; This represents the predicted stress distribution data; Represents pi; Represent the covariance matrix; Represents the natural exponential function; Represents a parameter vector (a column vector consisting of s, r, c, and t). Represents the mean vector; This represents the transpose operator.

[0074] It should be noted that the covariance matrix refers to the matrix that describes the covariance relationship between borehole spacing, borehole depth, charge amount, and detonation delay sequence in the joint probability distribution function.

[0075] S2.3: Extract the marginal probability distributions of borehole spacing, borehole depth, charge amount and detonation delay sequence parameters according to the joint probability distribution function. Calculate the cumulative distribution function of the marginal probability distributions of each parameter. Use the parameter value range where the cumulative probability reaches the preset confidence level as the search range of the corresponding parameter to complete the setting of the blasting parameter space.

[0076] Specifically, the marginal probability distributions of borehole spacing, borehole depth, charge amount, and detonation delay sequence are calculated by edge-mapping. The cumulative distribution function of the marginal probability distribution of each parameter is calculated, and the parameter value ranges where the cumulative probability reaches the preset confidence level are defined as the borehole spacing search range, borehole depth search range, charge amount search range, and detonation delay sequence search range, respectively, according to the engineering physical constraints and blasting design criteria of each parameter. The search ranges are combined to form a multi-dimensional blasting parameter space, thus completing the setting of the blasting parameter space.

[0077] It should be noted that engineering physical constraints and blasting design criteria refer to specific rules derived directly from the rock mass physical and mechanical parameters explicitly mentioned in the plan, the three-dimensional geometric data of the blasting face, and industry-standard specifications.

[0078] It should be noted that the search range for borehole spacing is a range of distance values ​​defined based on the requirements for the size of the broken rock mass and the effective range of the blasting stress wave. An exemplary range is 15 to 35 times the borehole diameter.

[0079] The borehole depth search range is the allowable range of vertical depth from the borehole opening to the bottom, defined based on the step height and the capabilities of the drilling equipment. An exemplary range is 70% to 110% of the step height.

[0080] The charge quantity search range is defined by the explosive unit consumption theory and vibration safety threshold, which defines the allowable weight range of explosives to be loaded in a single borehole. The range is determined by the product of explosive unit consumption and blasting volume.

[0081] The search range of the detonation delay sequence refers to the range of allowable values ​​for the detonation delay time between adjacent boreholes and rows, defined based on the vibration superposition effect and rock throwing control requirements. An exemplary range is typically from 15 milliseconds to 100 milliseconds per interval.

[0082] It should be noted that the pre-training of the underlying execution policy network is achieved by calling a random number generator to assign random decimal values ​​to the weight parameters of the underlying execution policy network. After the assignment, an initial weight matrix is ​​obtained. The feasibility of forward propagation of the network is verified through the initial weight matrix, the network infrastructure is initialized, and the pre-trained underlying execution policy network is obtained.

[0083] S2.4: After the setup is complete, embed the high-level meta-policy network and the low-level execution policy network into the meta-reinforcement learning agent and initialize them.

[0084] The high-level meta-policy network and the low-level execution policy network are connected within the meta-reinforcement learning agent framework. Random initial values ​​are assigned to the weight parameters of the high-level meta-policy network and the low-level execution policy network by calling a random number generator, thereby completing the weight parameter initialization. The high-level meta-policy network and the low-level execution policy network after weight parameter initialization constitute the initialized meta-reinforcement learning agent.

[0085] S2.5: Through the high-level meta-policy network in the meta-reinforcement learning intelligent body, feature extraction and target reasoning are performed on the predicted stress distribution and displacement field data to generate exploration sub-targets;

[0086] It should be noted that the pre-training process of the high-level meta-policy network is based on determining the dimension of the weight parameters according to the structure of the high-level meta-policy network, and obtaining the number of parameters that need to be initialized; calling a random number generator to generate random decimals that conform to the dimension to generate random weight values; directly assigning the random weight values ​​to the weight parameters of the high-level meta-policy network to form the initial weight matrix; the high-level meta-policy network with the initial parameter state after the assignment is completed, thus obtaining the pre-trained high-level meta-policy network.

[0087] Specifically, the high-level meta-policy network extracts spatial features from the predicted stress distribution and displacement field data through convolutional layers to obtain feature mapping data; it compresses the feature mapping data through pooling layers to generate dimensionality-reduced feature maps; it integrates the dimensionality-reduced feature maps through fully connected layers to generate low-dimensional feature vectors containing stress concentration risk and potential blasting efficiency; and it performs multi-objective collaborative reasoning based on the low-dimensional feature vectors. By analyzing the correlation between stress concentration risk and potential blasting efficiency in the low-dimensional feature vectors, it analyzes the balance between safety and efficiency and generates exploration sub-objectives.

[0088] S2.6 Based on the safety constraints in the exploration sub-targets, the underlying execution policy network in the meta-reinforcement learning intelligent body is used to sample the safety constraints in the blasting parameter space to generate a candidate blasting parameter combination set;

[0089] It should be noted that the safety constraints are obtained by analyzing the safety-related parameter dimensions in the sub-targets to obtain the quantitative requirements for the safety performance of the blasting scheme; by converting the quantitative requirements into numerical values, the peak particle velocity threshold is obtained, and the safety constraints with the threshold as the hard boundary are obtained.

[0090] Specifically, based on the safety constraints in the exploration sub-objectives, the underlying execution policy network within the meta-reinforcement learning intelligent body utilizes the safety constraints to directly generate borehole spacing values, borehole depth values, charge amount values, and detonation delay sequence values ​​that meet the safety constraints within the blasting parameter space, which is composed of the borehole spacing search range, borehole depth search range, charge amount search range, and detonation delay sequence search range. These values ​​are then combined to generate a candidate blasting parameter combination set.

[0091] S2.7: Predict the safety and efficiency results of candidate blasting parameter combinations using a meta-reinforcement learning agent, and obtain the predicted safety and efficiency indicators.

[0092] Specifically, the meta-reinforcement learning agent uses fixed physical information neural network physical constraint weight coefficients to perform dynamic simulation of the blasting scenario corresponding to each candidate blasting parameter combination in the candidate blasting parameter combination set, obtains the predicted peak mass velocity as a safety indicator, and determines the single-cycle blasting volume as an efficiency indicator based on the blasting fragmentation range.

[0093] S2.8: Calculate the reward value for each candidate blasting parameter combination in the candidate blasting parameter combination set based on the predicted safety and efficiency indicators, and select the candidate blasting parameter combination with the highest reward value as the current blasting parameter combination. The expression is:

[0094] ;

[0095] in, This represents the reward value for the current candidate combination of detonation parameters; The weighting coefficients representing efficiency indicators; Indicates the volume of a single-cycle blast; Indicates the maximum reference blasting volume; Indicates the weighting coefficient of safety indicators; This represents the predicted peak particle velocity; This indicates the preset vibration prediction safety threshold; This represents a linear penalty term for vibration velocities exceeding the vibration prediction safety threshold.

[0096] It should be noted that the weighting coefficient for efficiency indicators is a parameter used to balance the importance of efficiency indicators in the calculation of reward value. It is determined by the ratio of the mine's production target (such as the annual planned blasting volume) to the efficiency benchmark for a single blast. An exemplary value of 0.7 is used to ensure that the blasting plan prioritizes production needs without exceeding safety constraints. The weighting coefficient for safety indicators is a parameter used to reinforce the priority of safety indicators in the calculation of reward value. It is determined by the ratio of the maximum allowable peak mass velocity in the safety specifications to the actual blasting vibration risk. An exemplary value of 0.3 is used to ensure that the blasting... The vibration prediction value of the scheme is strictly lower than the vibration prediction safety threshold and retains a redundancy margin. The vibration prediction safety threshold is determined by obtaining the rock mass wave velocity and attenuation coefficient through geological exploration, combined with the surrounding building structure type, and the theoretical value is calculated using the Sadovsky formula based on the distance between the blasting point and the protected target. The theoretical value is then determined after being calibrated by the field vibration test data. An exemplary value range is 0.5~1.5 cm / s. Values ​​below 0.5 cm / s will excessively limit the blasting scale, resulting in a significant decrease in production efficiency, while values ​​above 1.5 cm / s will cause rock mass instability or vibration exceeding the standard risk.

[0097] The expression for calculating the theoretical value is:

[0098] ;

[0099] in, Indicates the safety threshold vibration velocity; , Indicates the site geological coefficient; Indicates the distance between the blast point and the protected target; Indicates the maximum charge per segment; Represents the structural sensitivity coefficient; Indicates the geological attenuation correction factor;

[0100] S2.9: Input the current combination of blasting parameters into the digital twin to update and optimize the high-level meta-strategy network parameters until the high-level meta-strategy network parameters reach the convergence condition, and then solidify the high-level meta-strategy network parameters into the blasting scheme evaluation strategy.

[0101] It should be noted that the convergence condition refers to the state when the change in the network parameters of the high-level meta-policy network is lower than the preset network parameter convergence threshold in consecutive iterations.

[0102] Specifically, the current blasting parameter combination is used in a digital twin. The digital twin generates corresponding safety and efficiency indicators based on the fixed physical information neural network physical constraint weight coefficients. The safety and efficiency indicators are used to guide the adjustment of the high-level meta-strategy network parameters, which are continuously optimized through iteration. When the change in the high-level meta-strategy network parameters in multiple consecutive iterations is lower than the preset network parameter convergence threshold, the high-level meta-strategy network parameters have reached the convergence condition. At this point, the high-level meta-strategy network parameters are saved and fixed as the blasting scheme evaluation strategy.

[0103] It should be noted that the network parameter convergence threshold is set based on the numerical stability requirements during the optimization of high-level meta-policy network parameters, and an exemplary value range is 1×10. -5 Up to 1×10 -3 less than 1×10 -5 This leads to a meaningless increase in training time, exceeding 1×10. -3 This can cause the network to stop before it converges, affecting the stability of the strategy.

[0104] S3. Call the preset target blasting scheme input element reinforcement learning agent, and perform multiple rounds of simulation deduction through the blasting scheme evaluation strategy to generate simulation deduction results.

[0105] S3.1: Input the preset target blasting scheme into the meta-reinforcement learning agent, analyze and extract features from the preset target blasting scheme, and generate scheme feature vectors;

[0106] It should be noted that the pre-set target blasting scheme refers to the blasting operation scheme that is planned in advance during the design stage of mine blasting engineering and includes specific parameter values. The blasting operation scheme consists of four parameters: borehole spacing, borehole depth, charge amount and detonation delay sequence. The parameter values ​​are determined based on geological exploration data, rock mechanics properties and engineering experience, and serve as the initial benchmark scheme for optimization and iteration of the meta-reinforcement learning agent.

[0107] Specifically, the parameters of borehole spacing, drilling depth, charge amount, and detonation delay sequence in the preset target blasting scheme are combined into a multi-dimensional input vector, which is then input into the high-level meta-policy network of the meta-reinforcement learning agent. The input layer of the high-level meta-policy network passes the multi-dimensional input vector to the hidden layer. The hidden layer extracts features through neuron weight calculation and activation function transformation to generate a high-dimensional feature map. At the same time, the output layer performs linear transformation and dimensionality reduction operations on the high-dimensional feature map to form the scheme feature vector.

[0108] S3.2: The feature vector of the blasting scheme is simulated and deduced in multiple rounds through the blasting scheme evaluation strategy to obtain the simulation and deduction instruction set;

[0109] It should be noted that parallel simulation refers to the process by which a digital twin simultaneously receives multiple sets of instructions from a simulation instruction set and performs synchronous dynamic simulation of the explosion scenario defined by each set of instructions based on fixed physical constraint weight coefficients.

[0110] Specifically, based on the scheme feature vector, multiple sets of simulation inference instructions are directly generated through the high-level meta-strategy network parameters fixed within the blasting scheme evaluation strategy. Each set of simulation inference instructions contains a specific combination of blasting parameter settings, and multiple sets of blasting parameter settings form a simulation inference instruction set.

[0111] It should be noted that the combination of blasting parameter settings refers to the specific values ​​of borehole spacing, borehole depth, charge amount, and detonation delay sequence explicitly specified in each simulation command, which are used to accurately describe the operating conditions of blasting operations.

[0112] S3.3: Parallel simulation and deduction of the simulation and deduction instruction set is performed using a digital twin, and the deduction indicators for each round are output;

[0113] Specifically, the simulation and deduction instruction set interacts with the digital twin. Based on the fixed physical information neural network physical constraint weight coefficients, the digital twin performs synchronous simulation and deduction of the instructions in the simulation and deduction instruction set, which contain multiple sets of specific blasting parameter settings. The simulation process reproduces the dynamic behavior of the blasting scene, thereby directly generating the peak mass velocity index and the single-cycle blasting volume index as the deduction index for each round.

[0114] It should be noted that the dynamic behavior of blasting scenarios refers to the dynamic changes in the interaction between the rock mass and the explosion energy during the blasting process, which originates from the physical information neural network solidified within the digital twin.

[0115] S3.4: The performance stability and robustness of each round of simulation indicators are evaluated and quantified by a multi-index fusion analysis algorithm to generate simulation results.

[0116] It should be noted that the multi-index fusion analysis algorithm refers to the algorithm that evaluates the stability and quantifies the robustness of each round of simulation inference, and generates simulation results through weighted fusion.

[0117] Specifically, the system receives the peak particle velocity index and single-cycle blast volume index generated from each round of simulation. Performance stability is assessed by statistically analyzing the fluctuations of each index value across multiple simulations; smaller fluctuations result in a higher stability score. The system then calculates the difference between the peak particle velocity and single-cycle blast volume index values ​​under the parameter disturbance scenario and their corresponding values ​​under the baseline scenario. The absolute value of the difference is used as the degree of deviation to obtain a robustness score; smaller deviations result in a higher robustness score. Finally, the stability score and robustness score are arithmetically averaged and fused to generate the simulation results.

[0118] S4. When the peak particle velocity in the simulation results exceeds the preset peak particle velocity threshold, parameter adjustment suggestions are generated.

[0119] S4.1: Extract the peak particle velocity from the simulation results. When the peak particle velocity exceeds the preset peak particle velocity threshold, trigger the parameter adjustment signal.

[0120] It should be noted that the preset peak particle velocity threshold is set by taking the minimum value calculated from the statutory benchmark value of the comprehensive blasting safety regulations, the theoretical value calculated based on the Sadovsky formula based on geological exploration data, and the environmental coefficient adjustment value of the distance to the protected target and the sensitivity of the structure. An exemplary range is 0.5 cm / s to 1.5 cm / s. A value below 0.5 cm / s will excessively limit blasting efficiency, while a value above 1.5 cm / s will cause structural damage risk.

[0121] Specifically, the peak particle velocity value is directly read from the simulation results and compared with the preset peak particle velocity threshold. If the peak particle velocity value is less than or equal to the preset peak particle velocity threshold, the current parameter configuration is maintained and the blasting scheme generation stage is directly entered. If the peak particle velocity value is greater than the preset peak particle velocity threshold, a parameter adjustment signal is immediately generated.

[0122] S4.2: Based on the parameter adjustment signal, initiate the correlation characteristic analysis of the current blasting parameter combination, and match the analysis results with the historical exploration data of the meta-reinforcement learning agent to generate matching results;

[0123] Specifically, the borehole spacing, drilling depth, charge amount, and detonation delay sequence in the current blasting parameter combination are combined into a multi-dimensional feature vector. Based on the Euclidean distance between the multi-dimensional feature vector and the feature vectors corresponding to each historical blasting parameter combination, a similarity value combination is obtained. From the similarity value combinations, the historical blasting parameter combination corresponding to the smallest Euclidean distance is selected, and the safety, efficiency, and stability dimensions of the historical blasting parameter combination recorded in the historical research dataset are read. The read historical dimension indicators are used as the matching result.

[0124] It should be noted that the historical exploration dataset refers to the dataset of explosive parameter combinations and their corresponding performance results explored by the meta-reinforcement learning agent in past training iterations.

[0125] S4.3: Through a comprehensive analysis mechanism, the matching results are evaluated and optimized in multiple dimensions, and parameter adjustment suggestions are generated.

[0126] Specifically, the comprehensive analysis mechanism initiates multi-dimensional evaluation. By comparing the security, efficiency, and stability indicators in the matching results with preset corresponding thresholds, it generates security, efficiency, and stability evaluation conclusions. The optimization decision follows the principle of prioritizing security. If security is not up to standard, parameters are adjusted first to reduce risk. If security is up to standard, parameters are optimized with stability as the next priority over efficiency. The optimization decision results are then directly converted into parameter adjustment suggestions.

[0127] It should be noted that optimization decision-making refers to the process of automatically determining the priority, direction, and specific magnitude of parameter adjustments in the design of blasting schemes, based on multi-dimensional evaluation conclusions of safety, efficiency, and stability, and following the principle of absolute priority for safety, secondly for stability, and lastly for efficiency.

[0128] It should be noted that the preset corresponding thresholds refer to the preset security dimension threshold, the preset efficiency dimension threshold, and the preset stability dimension threshold.

[0129] The preset safety threshold is set based on legal limits, the vibration propagation attenuation law under specific geological conditions of the mine, and the structure of surrounding buildings or facilities that need to be protected. An exemplary value range is 0.5cm / s to 2.5cm / s. Below 0.5cm / s, the blasting scheme will be too conservative, significantly reducing blasting production efficiency and increasing mining costs. Above 2.5cm / s, it will violate mandatory safety regulations and increase the risk of rock mass instability, slope slippage, or damage to surrounding building structures.

[0130] The preset efficiency threshold is set based on the mine's production plan targets, the theoretical operating capacity of drilling and loading equipment, and the minimum economic volume for a single blast. An exemplary value range is 70% to 120% of the planned blast volume, with 70% ensuring the economic feasibility of the blasting operation and 120% limited by the maximum processing capacity of the equipment.

[0131] The preset stability dimension threshold is based on statistical analysis of historical blasting data and multiple simulation results. It calculates the fluctuation range of peak mass velocity index and blasting volume index to measure the reliability of the parameter combination output results. An exemplary value range is a fluctuation range of 5% to 15%. 5% represents a highly stable and reliable output result, and 15% is an acceptable fluctuation threshold. Exceeding this threshold will indicate insufficient robustness of the result.

[0132] It should be noted that the principle of prioritizing safety means that in the process of parameter adjustment decision-making, the safety dimension of the blasting scheme is considered before the efficiency and stability dimensions.

[0133] S5. Feed back the parameter adjustment suggestions to the meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed the preset peak particle velocity threshold. Then generate a blasting plan and input the blasting plan into the gold mine production scheduling terminal.

[0134] S5.1: Input the parameter adjustment suggestions into the meta-reinforcement learning agent to update the high-level meta-policy network parameters in the meta-reinforcement learning agent, and obtain the updated high-level meta-policy network parameters.

[0135] Specifically, the meta-reinforcement learning agent reads the adjustment instructions for borehole spacing, borehole depth, charge quantity, and detonation delay sequence from the parameter adjustment suggestions; through the preset weight mapping rules within the meta-reinforcement learning agent, the adjustment instructions are directly converted into modifications to the corresponding weight values ​​in the high-level meta-policy network parameters; the modifications are applied to the current high-level meta-policy network parameters to generate new parameter values; and the new parameter values ​​are directly overwritten with the high-level meta-policy network parameters to obtain the updated high-level meta-policy network parameters.

[0136] It should be noted that the preset weight mapping rule refers to a set of logical criteria pre-set during the initialization phase of the meta-reinforcement learning agent, which is used to establish a direct correspondence between adjustment instructions and the modification amount of specific weight values ​​in the parameters of the high-level meta-policy network.

[0137] S5.2: Iteratively optimize the current blasting parameter combination based on the updated high-level meta-strategy network parameters until the peak particle velocity does not exceed the preset peak particle velocity threshold. Then, generate the optimized blasting parameter combination and input it as the blasting scheme into the gold mine production scheduling terminal.

[0138] Specifically, the meta-reinforcement learning agent receives the exploration sub-objectives output by the high-level meta-policy network through the underlying execution policy network and directly generates new blasting parameter combinations. These new blasting parameter combinations are then input into a digital twin for simulation to obtain the corresponding peak particle velocity values. The peak particle velocity values ​​are compared with a preset peak particle velocity threshold. If the peak particle velocity value exceeds the preset threshold, iterative optimization is performed again until the peak particle velocity value does not exceed the preset threshold, at which point the current blasting parameter combination is determined as the optimized blasting parameter combination. Finally, the optimized blasting parameter combination is sent to the gold mine production scheduling terminal as the blasting scheme.

[0139] This embodiment also provides a computer device applicable to the digital safety production management method for open-pit mines, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the digital safety production management method for open-pit mines as proposed in the above embodiment.

[0140] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0141] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the digital safety production management method for open-pit mines as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0142] In summary, this invention utilizes a meta-reinforcement learning agent to autonomously explore the blasting parameter space, enabling the efficient generation of blasting scheme evaluation strategies that combine adaptability and reliability. Through interactive learning between the agent and a digital twin, it achieves comprehensive performance evaluation of different combinations of blasting parameters under dynamic geological conditions, generating decision rules that balance dynamic control and fragmentation effects. This allows for rapid identification of parameter equilibrium points, improving the adaptability and reliability of blasting schemes, providing accurate decision-making basis for mine safety production, and ensuring a balance between the economic efficiency and safety of blasting effects.

[0143] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A digital safety production management method for open-pit mines, characterized in that: include, Collect the physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting operation face in the gold mining area, and use them as physical constraints to construct a digital twin; A blasting parameter space is set up in the digital twin, and the meta-reinforcement learning agent in the digital twin is initialized. The meta-reinforcement learning agent explores the safety and efficiency results corresponding to different combinations of blasting parameters in the blasting parameter space to obtain the blasting scheme evaluation strategy. The specific steps are as follows. By using a high-level meta-policy network within a meta-reinforcement learning intelligent body, feature extraction and target reasoning are performed on the predicted stress distribution and displacement field data to generate exploratory sub-targets. Based on the safety constraints in the exploration sub-objectives, the underlying execution policy network in the meta-reinforcement learning intelligent body is used to sample the safety constraints in the blasting parameter space to generate a candidate blasting parameter combination set. By using a meta-reinforcement learning agent to predict the safety and efficiency results of candidate blasting parameter combinations, the predicted safety and efficiency indicators are obtained. The reward value of each candidate blasting parameter combination in the candidate blasting parameter combination set is calculated based on the predicted safety and efficiency indicators, and the candidate blasting parameter combination with the largest reward value is taken as the current blasting parameter combination. Input the current combination of blasting parameters into the digital twin to update and optimize the parameters of the high-level meta-strategy network until the high-level meta-strategy network parameters reach the convergence condition, and then solidify the high-level meta-strategy network parameters into the blasting scheme evaluation strategy. The system calls upon a pre-set target blasting scheme input meta-reinforcement learning agent, performs multiple rounds of simulation and deduction through the blasting scheme evaluation strategy, and generates simulation and deduction results. When the peak particle velocity in the simulation results exceeds the preset peak particle velocity threshold, parameter adjustment suggestions are generated. The parameter adjustment suggestions are fed back to the meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed the preset peak particle velocity threshold, at which point a blasting scheme is generated.

2. The digital safety production management method for open-pit mines as described in claim 1, characterized in that: The process involves collecting physical and mechanical parameters of the rock mass and three-dimensional geometric data of the blasting face in the gold mine mining area, and using this data as physical constraints to construct a digital twin. The specific steps are as follows. The physical and mechanical parameters of the rock mass and the three-dimensional geometric morphology data of the blasting face are used as physical constraints and input into the physical information neural network. The encoder in the physical information neural network converts the three-dimensional geometric data of the blasting face into a latent space feature vector. The latent space feature vector and rock mass physical and mechanical parameters are fused and calculated using the decoder in the physical information neural network to generate predicted stress distribution and displacement field data. The predicted stress distribution and displacement field data are compared with the real-time monitored slope displacement and vibration data to obtain residual values. The physical constraint weight coefficients of the physical information neural network are then optimized based on the residual values. The optimized physical constraint weight coefficients are solidified into a digital twin.

3. The digital safety production management method for open-pit mines as described in claim 2, characterized in that: The blasting parameter space includes borehole spacing, borehole depth, charge amount, and detonation delay sequence.

4. The digital safety production management method for open-pit mines as described in claim 3, characterized in that: The specific steps for setting up the blasting parameter space in the digital twin and initializing the meta-reinforcement learning agent in the digital twin are as follows: Based on the predicted stress distribution and displacement field data, the joint probability distribution function of borehole spacing, borehole depth, charge amount and detonation delay sequence is calculated; Based on the joint probability distribution function, the marginal probability distributions of borehole spacing, borehole depth, charge amount, and detonation delay sequence parameters are extracted respectively; the cumulative distribution function of the marginal probability distribution of each parameter is calculated to obtain the high probability density interval as the search range of the corresponding parameter, thus completing the setting of the blasting parameter space; After the setup is complete, the high-level meta-policy network and the low-level execution policy network are embedded into the meta-reinforcement learning agent and initialized.

5. The digital safety production management method for open-pit mines as described in claim 1, characterized in that: The process involves calling a preset target blasting scheme input meta-reinforcement learning agent, performing multiple rounds of simulation deduction through a blasting scheme evaluation strategy, and generating simulation deduction results. The specific steps are as follows: The preset target blasting scheme is input into the meta-reinforcement learning agent, which analyzes and extracts features from the preset target blasting scheme to generate a scheme feature vector. By using a blasting scheme evaluation strategy, the scheme feature vectors are simulated and deduced in multiple rounds in parallel to obtain a simulation and deduction instruction set; Parallel simulation and deduction of the simulation and deduction instruction set are performed using a digital twin, and the simulation indicators for each round are output. The performance stability and robustness of each round of simulation indicators are evaluated and quantified by a multi-index fusion analysis algorithm, generating simulation results.

6. The digital safety production management method for open-pit mines as described in claim 5, characterized in that: When the peak particle velocity in the simulation results exceeds the preset peak particle velocity threshold, parameter adjustment suggestions are generated. The specific steps are as follows: Extract the peak particle velocity from the simulation results. When the peak particle velocity exceeds the preset peak particle velocity threshold, trigger the parameter adjustment signal. Based on the parameter adjustment signal, the correlation characteristic analysis of the current blasting parameter combination is initiated, and the analysis results are matched with the historical exploration data of the meta-reinforcement learning agent to generate matching results; The matching results are evaluated and optimized through a comprehensive analysis mechanism, generating parameter adjustment suggestions.

7. The digital safety production management method for open-pit mines as described in claim 6, characterized in that: The parameter adjustment suggestions are fed back to the meta-reinforcement learning agent for iterative optimization until the peak particle velocity does not exceed the preset peak particle velocity threshold, at which point a demolition plan is generated. The specific steps are as follows: Input the parameter adjustment suggestions into the meta-reinforcement learning agent to update the high-level meta-policy network parameters in the meta-reinforcement learning agent, and obtain the updated high-level meta-policy network parameters. The current blasting parameter combination is iteratively optimized based on the updated high-level meta-strategy network parameters until the peak particle velocity does not exceed the preset peak particle velocity threshold. Then, the optimized blasting parameter combination is generated and input as the blasting scheme to the gold mine production scheduling terminal.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the digital safety production management method for open-pit mines as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the digital safety production management method for open-pit mines as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tunnel deep-buried drainage ditch rapid blasting excavation method based on vertical drilling

    CN120627827A