Dynamic goaf three-zone grouting regulation and control method and system based on reinforcement learning
Through multi-source monitoring data analysis and digital twin technology based on reinforcement learning, the accuracy and applicability problems in the goaf three-belt grouting technology are solved, and efficient, stable and real-time decision-making support for grouting regulation are achieved.
Patent Information
- Application Number
- CN202510748346.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing three-belt grouting technology in goaf zones has problems such as insufficient accuracy, limited applicability, limited data processing capabilities and lack of dynamic optimization, resulting in unstable grouting effect.
Using reinforcement learning-based methods, a multi-source monitoring data analysis model and a dynamic grouting control model are constructed, combined with digital twin technology, real-time multi-source monitoring data analysis and dynamic control strategy generation are realized, and data processing and decision-making support are used for deep learning algorithms and reinforcement learning algorithms.
It improves the accuracy and efficiency of grouting control, avoids resource waste, enhances the applicability under complex geological conditions, realizes a comprehensive analysis of dynamic changes in space and time, provides real-time decision-making support, and improves the stability and reliability of grouting effect.
Smart Images

Figure CN120258341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of geotechnical engineering, and particularly relates to a dynamic regulation method and system for gob three-zone grouting based on reinforcement learning. Background Technique
[0002] A coal mine gob refers to the cavities or voids formed underground after coal resources are mined during the coal mining process. The existence of such areas will have a significant impact on mine safety, geological environment, and the lives of surrounding residents. The coal mine gob is an inevitable phenomenon during the coal mining process, and the resulting geological and safety problems need to be highly emphasized. The three zones of the gob (caving zone, fissure zone, and bending zone) refer to three typical strata movement areas formed during the coal mining process due to the deformation and failure of the overlying strata after the underground coal is mined out. Gob three-zone grouting refers to a technical means of injecting slurry (such as cement slurry, coal gangue slurry, etc.) into the caving zone, fissure zone, and bending zone during the treatment of coal mine gobs to fill the cavities, reinforce the strata, and improve the geological conditions. This method can effectively solve problems such as surface subsidence and mine water inrush caused by gobs, and is applicable to reinforcement scenarios in complex strata such as mines and tunnels.
[0003] In the field of gob three-zone grouting technology, although some technologies and methods have been applied, they still have many deficiencies, specifically as follows: 1) Lack of accuracy: Traditional grouting technologies usually rely on empirical judgment and it is difficult to achieve precise control of the grouting range. For example, when filling the fissures in the gob, due to inaccurate grouting range, some areas may not be effectively filled, or the slurry may flow into non-target areas, resulting in waste of resources; 2) Limited applicability: For complex geological conditions (such as high-stress and high-permeability areas), traditional grouting methods are difficult to meet the requirements of efficient filling, and the selection of grouting materials lacks pertinence; 3) Limited data processing ability: Most of the existing monitoring data are for static analysis, lacking the comprehensive analysis ability of spatio-temporal dynamic changes, and unable to provide real-time decision-making support for dynamic regulation of grouting; 4) Lack of dynamic optimization: Most of the existing grouting regulation strategies are static planning, unable to dynamically adjust grouting decisions (such as grouting volume, grouting position, grouting pressure) according to real-time monitoring data, resulting in unstable grouting effects. Summary of the Invention
[0004] In order to solve the problems of lack of accuracy, limited applicability, limited data processing ability, and lack of dynamic optimization existing in the prior art, the purpose of the present invention is to provide a dynamic regulation method and system for gob three-zone grouting based on reinforcement learning.
[0005] The technical solution adopted by the present invention is as follows: A dynamic regulation method for gob three-zone grouting based on reinforcement learning, comprising the following steps: Using a deep learning algorithm, construct a multi-source monitoring data analysis model, using a reinforcement learning algorithm, construct a grouting dynamic regulation model, and using digital twin technology, construct a gob three-zone digital twin model; According to the real-time multi-source monitoring data after spatio-temporal association, use the multi-source monitoring data analysis model to perform data analysis and obtain the real-time multi-source monitoring data analysis result; According to the real-time multi-source monitoring data analysis result, use the grouting dynamic regulation model to generate a grouting dynamic regulation strategy and obtain a real-time grouting dynamic regulation strategy; Use the gob three-zone digital twin model to visualize the real-time multi-source monitoring data analysis result and the real-time grouting dynamic regulation strategy, execute the real-time grouting dynamic regulation strategy, and continue the data acquisition step.
[0006] Furthermore, using a deep learning algorithm, construct a multi-source monitoring data analysis model, using a reinforcement learning algorithm, construct a grouting dynamic regulation model, and using digital twin technology, construct a gob three-zone digital twin model, including the following steps: Using a deep learning algorithm, construct an initial multi-source monitoring data analysis model, using a reinforcement learning algorithm, construct an initial grouting dynamic regulation model; Construct a gob three-zone three-dimensional simulation model, and combine the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model, and use digital twin technology to construct a gob three-zone digital twin model; Collect a number of historical multi-source monitoring data of the gob three-zone in the physical world corresponding to the gob three-zone digital twin model in the digital world, and perform preprocessing to obtain a number of preprocessed historical multi-source monitoring data; According to the gob three-zone three-dimensional simulation model and a number of preprocessed historical multi-source monitoring data, train the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model of the gob three-zone digital twin model to obtain a final multi-source monitoring data analysis model and a final grouting dynamic regulation model.
[0007] Furthermore, the multi-source monitoring data analysis model is constructed based on the 3D-DBN-CNN-LSTM-ST-CNN algorithm, and the multi-source monitoring data analysis model includes a three-dimensional space feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm. The three-dimensional space feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module; The grouting dynamic regulation model is constructed based on the MPO-MOGRPO-ICPO algorithm. The grouting dynamic regulation model includes a meta-strategy optimization module constructed based on the MPO algorithm, a regulation strategy generation module constructed based on the MOGRPO algorithm, and a regulation strategy optimization module constructed based on the ICPO algorithm, which are connected in sequence. The regulation strategy generation module includes a set of objective functions, an experience replay pool, an Actor network, and an agent. The agent is respectively connected to the set of objective functions, the experience replay pool, and the Actor network. The Actor network is connected to the regulation strategy optimization module.
[0008] Furthermore, using the deep learning algorithm, an initial multi-source monitoring data analysis model is constructed, and using the reinforcement learning algorithm, an initial grouting dynamic regulation model is constructed, including the following steps: Using the 3D-DBN-CNN-LSTM-ST-CNN algorithm, an initial multi-source monitoring data analysis model is constructed; Using the MPO-MOGRPO-ICPO algorithm, an initial grouting dynamic regulation model is constructed; the initial grouting dynamic regulation model includes an initial meta-strategy optimization module, an initial regulation strategy generation module, and an initial regulation strategy optimization module; A set of objective functions, an experience replay pool, an Actor network, and an agent are set for the initial regulation strategy generation module, and the initial network parameters of the Actor network are used as the output parameters of the initial meta-strategy optimization module; Taking the grouting dynamic regulation strategy generation problem as the simulation environment of the initial regulation strategy generation module, and setting the action space and state space for the agent of the initial regulation strategy generation module; Taking the minimization of the strategy generation error as the optimization goal, defining the fitness function of the initial regulation strategy optimization module, and using the output of the Actor network of the initial regulation strategy generation module as the input of the initial regulation strategy optimization module.
[0009] Furthermore, according to the three-zone three-dimensional simulation model of the goaf and a number of preprocessed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model of the three-zone digital twin model of the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic regulation model, including the following steps: Combining the three-zone three-dimensional simulation model of the goaf with a number of preprocessed historical multi-source monitoring data, and dividing them into a model training set and a model test set according to a ratio of 7:3; According to the model training set, the initial multi-source monitoring data analysis model of the three-zone digital twin model of the goaf is trained to obtain an optimized multi-source monitoring data analysis model, and a number of historical multi-source monitoring data analysis results are generated; Test the optimized multi-source monitoring data analysis model according to the model test set. If the test accuracy rate is greater than the accuracy threshold, output the final multi-source monitoring data analysis model; otherwise, continue training. Traverse all objective functions in the objective function set, and train the initial grouting dynamic regulation model based on several historical multi-source monitoring data analysis results to obtain the final grouting dynamic regulation model, and store the historical grouting dynamic regulation experience generated during the training process in the experience replay pool.
[0010] Furthermore, the historical multi-source monitoring data includes historical strata monitoring data, historical environmental monitoring data, and historical equipment operation monitoring data of the three zones in the goaf. The real-time multi-source monitoring data includes real-time strata monitoring data, real-time environmental monitoring data, and real-time equipment operation monitoring data of the three zones in the goaf.
[0011] Furthermore, based on the digital twin model of the three zones in the goaf, collect the real-time multi-source monitoring data of the three zones in the goaf, and perform spatio-temporal association on the real-time multi-source monitoring data to obtain the spatio-temporally associated real-time multi-source monitoring data, including the following steps: Collect the real-time multi-source monitoring data of the three zones in the goaf in the physical world, perform preprocessing, and input the preprocessed real-time multi-source monitoring data into the digital twin model of the three zones in the goaf. Align the preprocessed real-time multi-source monitoring data in terms of time stamps to obtain the time-aligned real-time multi-source monitoring data. Perform spatial registration on the time-aligned real-time multi-source monitoring data to obtain the spatially registered real-time multi-source monitoring data. Extract the real-time time features and real-time spatial features of the spatially registered real-time multi-source monitoring data, and perform feature fusion on the real-time time features and real-time spatial features to obtain the real-time fused spatio-temporal features. Perform spatio-temporal association on the real-time fused spatio-temporal features in the digital twin model of the three zones in the goaf to obtain the spatio-temporally associated real-time multi-source monitoring data.
[0012] Furthermore, according to the spatio-temporally associated real-time multi-source monitoring data, use the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis results, including the following steps: Use the three-dimensional spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time three-dimensional spatial features of the three-dimensional simulation model of the three zones in the goaf in the digital twin model of the three zones in the goaf. Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the spatio-temporally associated real-time multi-source monitoring data. The structured data feature extraction module using the multi-source monitoring data analysis model extracts the real-time structured data features of the real-time multi-source monitoring data after spatio-temporal association; According to the real-time three-dimensional space features, real-time unstructured data features, and real-time structured data features, the multi-source monitoring data analysis module of the multi-source monitoring data analysis model is used to perform data analysis to obtain the real-time multi-source monitoring data analysis results.
[0013] Furthermore, according to the real-time multi-source monitoring data analysis results, the grouting dynamic regulation model is used to generate the grouting dynamic regulation strategy to obtain the real-time grouting dynamic regulation strategy, including the following steps: According to the real-time multi-source monitoring data analysis results, the meta-strategy optimization module of the grouting dynamic regulation model adjusts the Actor network of the regulation strategy generation module of the grouting dynamic regulation model to obtain the adjusted Actor network; Select the real-time objective function from the set of objective functions of the grouting dynamic regulation model, and randomly extract a number of historical grouting dynamic regulation experiences from the experience replay pool of the grouting dynamic regulation model based on the real-time objective function; According to the real-time multi-source monitoring data analysis results, the state space of the intelligent agent is adjusted to obtain the adjusted state space, and according to a number of historical grouting dynamic regulation experiences, the action space of the intelligent agent is adjusted to obtain the adjusted action space; Based on the real-time objective function, the adjusted state space, and the adjusted action space, the intelligent agent of the regulation strategy generation module is used to control the adjusted Actor network to generate the real-time grouting dynamic regulation action probability distribution; The regulation strategy optimization module of the grouting dynamic regulation model is used to optimize the real-time grouting dynamic regulation action probability distribution to obtain the real-time grouting dynamic regulation strategy.
[0014] A goaf three-zone grouting dynamic regulation system based on reinforcement learning is used to implement the goaf three-zone grouting dynamic regulation method, including a model construction unit, a spatio-temporal association unit, a data analysis unit, a regulation strategy generation unit, and a visualization unit connected in sequence.
[0015] The beneficial effects of the present invention are: A gob three-zone grouting dynamic regulation method and system based on reinforcement learning provided by the present invention, and a multi-source monitoring data analysis model constructed, avoid relying on empirical judgment, comprehensively consider the information of strata, environment and equipment operation included in multi-source monitoring data, and combine three-dimensional space information such as the geometric shape, spatial distribution and volume change of the gob three zones, improve the accuracy and efficiency of grouting regulation, and avoid waste of resources; can comprehensively analyze multi-source monitoring data under complex geological conditions, provide a scientific basis for the selection of grouting materials and the formulation of grouting strategies, and at the same time, can dynamically adjust grouting parameters according to different geological conditions, improve the applicability under different geological environments; combine deep learning algorithms and reinforcement learning algorithms, realize automated multi-source monitoring data analysis, improve data processing efficiency and accuracy, and have the comprehensive analysis ability of spatio-temporal dynamic changes, provide real-time decision support for grouting dynamic regulation; the constructed grouting dynamic regulation model adopts a dynamic optimization mechanism, can adjust grouting decisions in a timely manner according to the results of real-time multi-source monitoring data analysis, and improve the stability and reliability of grouting effects.
[0016] Other beneficial effects of the present invention will be further described in the specific implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flow block diagram of the gob three-zone grouting dynamic regulation method based on reinforcement learning in the present invention.
[0018] Figure 2 is a structural block diagram of the gob three-zone grouting dynamic regulation system based on reinforcement learning in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.
[0020] Embodiment 1: As Figure 1 shown, this embodiment provides a gob three-zone grouting dynamic regulation method based on reinforcement learning, including the following steps: S1: Use deep learning algorithms to construct a multi-source monitoring data analysis model, use reinforcement learning algorithms to construct a grouting dynamic regulation model, and use digital twin technology to construct a gob three-zone digital twin model, including the following steps: S1-1: Use deep learning algorithms to construct an initial multi-source monitoring data analysis model, and use reinforcement learning algorithms to construct an initial grouting dynamic regulation model; The multi-source monitoring data analysis model is constructed based on the 3D-Deep Belief Network (DBN)-Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM)-Spatio-Temporal Convolutional Neural Network (ST-CNN) algorithm. The multi-source monitoring data analysis model includes a three-dimensional spatial feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm. The three-dimensional spatial feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module; The 3D-DBN of the three-dimensional spatial feature extraction module is a deep learning algorithm that can extract complex spatial features from three-dimensional simulation models (such as three-dimensional data, three-dimensional meshes, or voxel data), improve the extraction accuracy of three-dimensional spatial features, and provide high-quality three-dimensional features for subsequent multi-source data analysis. Through deep learning technology, it can process complex spatial distribution and structural information and enhance the representational ability of the model. The CNN of the unstructured data feature extraction module is good at processing unstructured data such as images and videos, can extract local features and hierarchical information therein, improve the parsing ability of unstructured data (such as images, videos, etc.), effectively capture key information therein, and through the hierarchical feature extraction of the CNN, can better identify patterns in the data and provide support for subsequent analysis. The LSTM of the structured data feature extraction module is a variant of the recurrent neural network, good at processing time series data, can capture long-term dependencies in the data, improve the processing ability of time series data, especially perform excellently when the data has long-term dependencies, and through the temporal feature extraction of the LSTM, can better understand the dynamic changes of the data and provide support for data analysis. The ST-CNN of the multi-source monitoring data analysis module is a deep learning algorithm that combines spatial and temporal information, can process spatio-temporal data and extract its features, and through spatio-temporal feature fusion, can comprehensively analyze the dynamic changes of the three zones in the goaf, improve the accuracy and real-time performance of data analysis, and provide reliable data support for tasks such as grouting dynamic regulation; The grouting dynamic regulation model is constructed based on the Meta-Policy Optimization (MPO)-Multi-Objective Group Relative Policy Optimization (MOGRPO)-Improved Crested Porcupine Optimizer (ICPO) algorithm. The grouting dynamic regulation model includes a meta-policy optimization module constructed based on the MPO algorithm, a regulation strategy generation module constructed based on the MOGRPO algorithm, and a regulation strategy optimization module constructed based on the ICPO algorithm, which are connected in sequence. The regulation strategy generation module includes an objective function set, an experience replay pool, an Actor network, and an agent. The agent is respectively connected to the objective function set, the experience replay pool, and the Actor network. The Actor network is connected to the regulation strategy optimization module; The meta-policy optimization module is used to regulate the initial network parameters of the Actor network in the policy generation module, so that these parameters can quickly adapt to the new and unseen multi-source monitoring data analysis results, improve the generalization ability of the model, and even update the Actor network based on previous learning experience under the unseen multi-source monitoring data analysis results, improving the adaptability of the regulation policy generation model; the objective function set of the regulation policy generation module can handle multiple conflicting optimization objectives, such as equipment operation efficiency, parameter regulation cost, operation safety, etc., and generate regulation policies that balance these objectives. The agent learns historical experience through the experience replay pool and continuously optimizes its own policy generation ability. The agent controls the Actor network according to the learned experience to generate more effective grouting dynamic regulation policies. The design of the experience replay pool and the agent enables the model to continuously learn and optimize, improving the quality of policy generation. Since the regulation policy generation module adopts the method of swarm exploration, it can avoid falling into local optimal solutions to a certain extent. The Actor network outputs the distribution probability of actions in a given state, and the goal is to learn an optimal policy, that is, to maximize the long-term cumulative reward. In the continuous action space, the Actor network usually outputs a mean value and optional variance parameters to describe the probability distribution of actions. The Critic network is responsible for evaluating the value of a given state, that is, predicting the expected return that can be obtained starting from this state and following the current policy, and usually outputs a scalar value representing the value of the state or the state-action value. The experience replay pool is used to store historical experience for reuse during the training process. The regulation policy generation module directly updates the Actor network through gradients, eliminating the Critic network in traditional reinforcement learning, making the algorithm structure simpler; the ICPO algorithm of the regulation policy optimization module is an efficient optimization algorithm that can quickly find the optimal or approximate optimal solution. By optimizing the distribution probability of actions in a given state output by the Actor network, it further improves the accuracy and generation efficiency of the real-time grouting dynamic regulation policy; Using deep learning algorithms, construct an initial multi-source monitoring data analysis model, and using reinforcement learning algorithms, construct an initial grouting dynamic regulation model, including the following steps: S1-1-1: Use the 3D-DBN-CNN-LSTM-ST-CNN algorithm to construct an initial multi-source monitoring data analysis model; S1-1-2: Use the MPO-MOGRPO-ICPO algorithm to construct an initial grouting dynamic regulation model; the initial grouting dynamic regulation model includes an initial meta-policy optimization module, an initial regulation policy generation module, and an initial regulation policy optimization module; S1-1-3: Set up a set of objective functions, an experience replay pool, an Actor network, and an agent for the initial control strategy generation module, and use the initial network parameters of the Actor network as the output parameters of the initial meta-policy optimization module; S1-1-4: Take the problem of grouting dynamic control strategy generation as the simulation environment of the initial control strategy generation module, and set the action space and state space for the agent of the initial control strategy generation module; S1-1-5: Define the fitness function of the initial control strategy optimization module with the goal of minimizing the strategy generation error, and use the output of the Actor network of the initial control strategy generation module as the input of the initial control strategy optimization module; S1-2: Construct a three-dimensional simulation model of the three zones in the goaf, and combine the initial multi-source monitoring data analysis model and the initial grouting dynamic control model. Use digital twin technology to construct a digital twin model of the three zones in the goaf, including the following steps: S1-2-1: Collect the geological data, Geographic Information System (GIS) data, and three-dimensional scan data of the three zones in the goaf, and use three-dimensional simulation technology to construct a three-dimensional simulation model of the three zones in the goaf based on the GIS data and three-dimensional scan data; S1-2-2: Use digital twin technology to perform digital mapping on the three-dimensional simulation model of the three zones in the goaf to obtain the digital twin framework of the three zones in the goaf; S1-2-3: Set the initial multi-source monitoring data analysis model and the initial grouting dynamic control model in the digital twin framework of the three zones in the goaf; S1-2-4: Convert the real-time multi-source monitoring data into real-time data streams, and connect the real-time data streams to the initial multi-source monitoring data analysis model and the initial grouting dynamic control model in the digital twin framework of the three zones in the goaf to obtain a three-dimensional simulation model of the three zones in the goaf; S1-3: Collect a number of historical multi-source monitoring data of the three zones in the goaf in the physical world corresponding to the digital twin model of the three zones in the digital world, and perform preprocessing to obtain a number of preprocessed historical multi-source monitoring data; The historical multi-source monitoring data includes historical strata monitoring data, historical environmental monitoring data, and historical equipment operation monitoring data of the three zones in the goaf; The three zones in the goaf include the caving zone, the fissure zone, and the bending zone; The strata monitoring data includes: the caving height of the strata in the caving zone (the maximum height of the caving zone and its change over time), the distribution and particle size of the caved rock blocks (monitoring the distribution and particle size of the caved rock blocks through borehole camera or ground penetrating radar), the movement speed of the strata (recording the speed and dynamic changes during the caving process of the strata through monitoring equipment), and the caving range (the horizontal extension range of the caving zone); The fracture development height of the fracture zone (the maximum height of the fracture zone and its change over time), the fracture width and density (detecting the width and density of the fractures through borehole camera or ultrasonic detection), the fracture direction (the strike and dip of the fractures, used to evaluate the stability of the strata), and the water seepage situation (the water seepage volume and water flow direction in the fracture zone, used to judge whether it may cause groundwater pollution); The bending deformation amount of the bending zone (monitoring the bending degree of the strata through strain gauges or inclinometers), the thickness of the bending zone (the thickness of the bending zone from the top of the fracture zone to the unaffected strata), the strata displacement (the horizontal and vertical displacements of the strata in the bending zone), and the stress distribution (monitoring the stress changes in the bending zone through stress gauges); The environmental monitoring data includes the environmental temperature, environmental humidity, environmental altitude, and environmental longitude and latitude of the caving zone, fracture zone, and bending zone; The equipment operation monitoring data includes the operation power, operation current, operation voltage, operation mode, operation parameters, and fault codes of the grouting equipment; The preprocessing includes data cleaning, denoising, normalization, etc., to ensure the data quality and provide data support for subsequent model training and optimization; S1-4: According to the three-dimensional simulation model of the three zones in the goaf and a number of preprocessed historical multi-source monitoring data, train the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model of the digital twin model of the three zones in the goaf to obtain the final multi-source monitoring data analysis model and the final grouting dynamic regulation model, including the following steps: S1-4-1: Combine the three-dimensional simulation model of the three zones in the goaf with a number of preprocessed historical multi-source monitoring data, and divide them into a model training set and a model test set according to a ratio of 7:3; S1-4-2: According to the model training set, train the initial multi-source monitoring data analysis model of the digital twin model of the three zones in the goaf to obtain an optimized multi-source monitoring data analysis model, and generate a number of historical multi-source monitoring data analysis results; S1-4-3: According to the model test set, test the optimized multi-source monitoring data analysis model. If the test accuracy rate is greater than the accuracy rate threshold, output the final multi-source monitoring data analysis model; otherwise, continue training; S1-4-4: Traverse all the objective functions in the objective function set, train the initial grouting dynamic regulation model based on the analysis results of several historical multi-source monitoring data, obtain the final grouting dynamic regulation model, and store the historical grouting dynamic regulation experience generated during the training process in the experience replay pool; S2: Based on the digital twin model of the three zones in the goaf, collect the real-time multi-source monitoring data of the three zones in the goaf, and perform spatio-temporal association on the real-time multi-source monitoring data to obtain the spatio-temporally associated real-time multi-source monitoring data, including the following steps: S2-1: Collect the real-time multi-source monitoring data of the three zones in the goaf in the physical world, perform preprocessing, and input the preprocessed real-time multi-source monitoring data into the digital twin model of the three zones in the goaf; The real-time multi-source monitoring data includes the real-time strata monitoring data, real-time environmental monitoring data, and real-time equipment operation monitoring data of the three zones in the goaf; S2-2: Align the preprocessed real-time multi-source monitoring data at the time stamp to obtain the time-aligned real-time multi-source monitoring data; S2-3: Perform spatial registration on the time-aligned real-time multi-source monitoring data to obtain the spatially registered real-time multi-source monitoring data; S2-4: Extract the real-time time features and real-time spatial features of the spatially registered real-time multi-source monitoring data, and perform feature fusion on the real-time time features and real-time spatial features to obtain the real-time fused spatio-temporal features; S2-5: Perform spatio-temporal association on the real-time fused spatio-temporal features in the digital twin model of the three zones in the goaf to obtain the spatio-temporally associated real-time multi-source monitoring data; S3: According to the spatio-temporally associated real-time multi-source monitoring data, use the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis results, including the following steps: S3-1: Use the three-dimensional spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time three-dimensional spatial features of the three-dimensional simulation model of the three zones in the goaf in the digital twin model of the three zones in the goaf; S3-2: Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the spatio-temporally associated real-time multi-source monitoring data; S3-3: Use the structured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time structured data features of the spatio-temporally associated real-time multi-source monitoring data; S3-4: According to the real-time three-dimensional spatial features, real-time unstructured data features, and real-time structured data features, use the multi-source monitoring data analysis module of the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis results; The analysis results of real-time multi-source monitoring data include the analysis results of real-time grouting equipment anomalies, real-time grouting demand analysis results, real-time grouting position analysis results, etc.; S4: According to the analysis results of real-time multi-source monitoring data, use the grouting dynamic regulation model to generate a grouting dynamic regulation strategy, and obtain the real-time grouting dynamic regulation strategy, including the following steps: S4-1: According to the analysis results of real-time multi-source monitoring data, use the meta-strategy optimization module of the grouting dynamic regulation model to adjust the Actor network of the regulation strategy generation module of the grouting dynamic regulation model, and obtain the adjusted Actor network; S4-2: Select a real-time objective function from the set of objective functions of the grouting dynamic regulation model, and randomly extract a number of historical grouting dynamic regulation experiences from the experience replay pool of the grouting dynamic regulation model based on the real-time objective function; S4-3: According to the analysis results of real-time multi-source monitoring data, adjust the state space of the intelligent agent to obtain the adjusted state space, and adjust the action space of the intelligent agent according to a number of historical grouting dynamic regulation experiences to obtain the adjusted action space; S4-4: Based on the real-time objective function, the adjusted state space, and the adjusted action space, use the intelligent agent of the regulation strategy generation module to control the adjusted Actor network to generate the real-time grouting dynamic regulation action probability distribution, including the following steps: S4-4-1: Based on the real-time objective function, use the intelligent agent of the regulation strategy generation module to control the adjusted Actor network to generate the probability distribution of all possible actions in the adjusted action space corresponding to each real-time state in the adjusted state space; S4-4-2: Integrate the probability distributions of all possible actions of the same real-time state to obtain the real-time probability distribution sequence; S4-4-2: Traverse all real-time states in the adjusted state space to obtain the real-time grouting dynamic regulation action probability distribution composed of a number of real-time probability distribution sequences; S4-5: Use the regulation strategy optimization module of the grouting dynamic regulation model to optimize the real-time grouting dynamic regulation action probability distribution to obtain the real-time grouting dynamic regulation strategy, including the following steps: S4-5-1: Encode the real-time optimization parameter values of the real-time grouting dynamic regulation action probability distribution into the solution vector of the regulation strategy optimization module, and set the ICPO population parameters and the maximum number of iterations; S4-5-2: Initialize according to the solution vector and the ICPO population parameters to obtain a number of initial solutions; the initial solutions correspond to the initial real-time optimization parameter values; The formula is:
[0021] In the formula, is the initial ICPO individual of the Circle chaotic map, i.e., the initial solution; is the randomly generated initial ICPO individual; is the ICPO individual indicator; Compared with the population randomly distributed by the Circle chaotic map sequence to generate the initial population, the initial position distribution of the improved ICPO individuals is more uniform, expanding the search range of the algorithm in space, increasing the diversity of the population positions, and improving to a certain extent the defect that the algorithm is prone to fall into local extrema, thereby improving the optimization efficiency of the algorithm; S4-5-3: Introduce a cyclic population reduction mechanism to limit the number of individuals in the ICPO population parameters to obtain the updated ICPO population parameters for the next iteration; The formula is:
[0022] In the formula, is the number of individuals in the ICPO population parameters at the -th iteration; is the number of individuals in the ICPO population parameters at the -th iteration; is the minimum value of the number of individuals in the ICPO population parameters; is the function evaluation parameter; is the function evaluation loop parameter; is the maximum function evaluation loop parameter; t is the iteration number indicator; S4-5-4: Calculate the initial fitness value of the initial ICPO individuals in the initial ICPO population according to the fitness function; The formula for the fitness value is:
[0023] In the formula, is the fitness function; is the mean square error function, used to obtain the policy generation error; is the true output value of the -th sample; is the ideal output value of the -th sample; is the sample indicator; N is the total number of samples; S4-5-5: Update the initial ICPO population using the first defense strategy, the second defense strategy, the third defense strategy, and the fourth defense strategy according to the initial fitness value and the updated ICPO population parameters to obtain the updated ICPO population; The formula for the first defense strategy is:
[0024] Wherein, is the updated ICPO individual within the first defense range; is the initial ICPO individual within the first defense range; is a random number based on the normal distribution; is a random value within the interval [0, 1]; is the optimal solution within the first defense range; is a vector generated between the true optimal solution and the randomly selected optimal solution from the ICPO population within the first defense range; is the ICPO individual indicator; is the iteration indicator; The formula for the second defense strategy is:
[0025] Wherein, is the updated ICPO individual within the second defense range; is the initial ICPO individual within the second defense range; is the search upper limit vector of the second defense range; is a random value within the interval [0, 1]; are respectively the th initial ICPO individuals; are both two random integers between [1, ; is a vector generated between the true optimal solution and the randomly selected optimal solution from the ICPO population within the second defense range; The formula for the third defense strategy is:
[0026] Wherein, is the updated ICPO individual within the third defense range; is the initial ICPO individual within the third defense range; is the search upper limit vector of the third defense range; are respectively the th initial ICPO individuals; is a random integer between [1, ; is the odor diffusion factor defined by the fitness function; is the defense factor; is the search direction control parameter; The formula for the fourth defense strategy is:
[0027] In the formula, is the updated ICPO individual within the fourth defense range; is the initial ICPO individual within the fourth defense range; is the optimal solution within the fourth defense range; are all random values within the interval [0, 1]; is the defense factor; is the search direction control parameter; is the average force affecting the search direction; is the convergence speed factor; S4-5-6: Use the dynamic reverse learning algorithm to perform dynamic reverse learning on the updated ICPO population to generate a dynamically reversed ICPO population; The formula is:
[0028] In the formula, is the dynamically reversed ICPO individual; is the decreasing inertia coefficient; are the maximum and minimum values of the vector space respectively; is the updated ICPO individual; S4-5-7: According to the fitness function, calculate the fitness values of all ICPO individuals in the updated ICPO population and the dynamically reversed ICPO population, take the ICPO individual with the minimum fitness value as the optimal individual, and retain the optimal individual; S4-5-8: If the number of iterations of iterative optimization reaches the maximum number of iterations or the fitness value of the optimal individual meets the requirements, output the optimal solution corresponding to the optimal individual, and decode the solution vector of the optimal solution to obtain the optimal real-time optimization parameter value; S4-5-9: According to the optimal real-time optimization parameter value, optimize the real-time grouting dynamic regulation action probability distribution to obtain the real-time grouting dynamic regulation strategy; The real-time grouting dynamic regulation strategy includes decisions on adjusting the real-time operating parameters of grouting equipment, real-time grouting pressure control decisions, real-time grouting type control decisions, real-time grouting volume control decisions, etc.; S5: Use the digital twin model of the three zones in the goaf to visualize the results of real-time multi-source monitoring data analysis and the real-time grouting dynamic regulation strategy, execute the real-time grouting dynamic regulation strategy, and continue the data acquisition step.
[0029] Example 2: Such as Figure 2As shown in the figure, this embodiment provides a gob three-zone grouting dynamic regulation system based on reinforcement learning, which is used to implement the gob three-zone grouting dynamic regulation method, including a model construction unit, a spatio-temporal correlation unit, a data analysis unit, a regulation strategy generation unit, and a visualization unit connected in sequence.
[0030] The model construction unit is used to construct a multi-source monitoring data analysis model using a deep learning algorithm, construct a grouting dynamic regulation model using a reinforcement learning algorithm, and construct a gob three-zone digital twin model using digital twin technology; The spatio-temporal correlation unit is used to collect real-time multi-source monitoring data of the gob three-zone based on the gob three-zone digital twin model, and perform spatio-temporal correlation on the real-time multi-source monitoring data to obtain the real-time multi-source monitoring data after spatio-temporal correlation; The data analysis unit is used to perform data analysis on the real-time multi-source monitoring data after spatio-temporal correlation using the multi-source monitoring data analysis model to obtain the real-time multi-source monitoring data analysis result; The regulation strategy generation unit is used to generate a grouting dynamic regulation strategy based on the real-time multi-source monitoring data analysis result using the grouting dynamic regulation model to obtain the real-time grouting dynamic regulation strategy; The visualization unit is used to visualize the real-time multi-source monitoring data analysis result and the real-time grouting dynamic regulation strategy using the gob three-zone digital twin model, execute the real-time grouting dynamic regulation strategy, and continue the data collection step.
[0031] The gob three-zone grouting dynamic regulation method and system provided by the present invention construct a multi-source monitoring data analysis model, which avoids relying on empirical judgment, comprehensively considers the information of the formation, environment, and equipment operation included in the multi-source monitoring data, and combines the three-dimensional spatial information such as the geometric shape, spatial distribution, and volume change of the gob three-zone, improving the accuracy and efficiency of grouting regulation and avoiding resource waste; it can comprehensively analyze multi-source monitoring data under complex geological conditions, provide a scientific basis for the selection of grouting materials and the formulation of grouting strategies, and at the same time, can dynamically adjust grouting parameters according to different geological conditions, improving the applicability under different geological environments; it combines deep learning algorithms and reinforcement learning algorithms to realize automated multi-source monitoring data analysis, improving data processing efficiency and accuracy, and having the comprehensive analysis ability of spatio-temporal dynamic changes, providing real-time decision support for grouting dynamic regulation; the constructed grouting dynamic regulation model adopts a dynamic optimization mechanism, which can adjust grouting decisions in a timely manner according to the real-time multi-source monitoring data analysis result, improving the stability and reliability of grouting effects.
[0032] The present invention is not limited to the above optional embodiments, and any person can obtain other various forms of products under the inspiration of the present invention. The above specific embodiments should not be construed as limiting the protection scope of the present invention, and the protection scope of the present invention should be defined by the claims, and the specification can be used to interpret the claims.
Claims
1. A dynamic regulation method for gob three-zone grouting based on reinforcement learning, characterized in that: It includes the following steps: Using deep learning algorithms, construct a multi-source monitoring data analysis model, using reinforcement learning algorithms, construct a grouting dynamic regulation model, and using digital twin technology, construct a digital twin model of the three zones in the goaf; Based on the digital twin model of the three zones in the goaf, collect real-time multi-source monitoring data of the three zones in the goaf, and perform spatio-temporal association on the real-time multi-source monitoring data to obtain the real-time multi-source monitoring data after spatio-temporal association; According to the real-time multi-source monitoring data after spatio-temporal association, use the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis results; According to the real-time multi-source monitoring data analysis results, use the grouting dynamic regulation model to generate a grouting dynamic regulation strategy to obtain the real-time grouting dynamic regulation strategy; Use the digital twin model of the three zones in the goaf to visualize the real-time multi-source monitoring data analysis results and the real-time grouting dynamic regulation strategy, execute the real-time grouting dynamic regulation strategy, and continue the data collection step.
2. The gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 1, wherein: Using deep learning algorithms, construct a multi-source monitoring data analysis model, using reinforcement learning algorithms, construct a grouting dynamic regulation model, and using digital twin technology, construct a digital twin model of the three zones in the goaf, including the following steps: Using deep learning algorithms, construct an initial multi-source monitoring data analysis model, using reinforcement learning algorithms, construct an initial grouting dynamic regulation model; Construct a three-dimensional simulation model of the three zones in the goaf, and combine the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model, and use digital twin technology to construct a digital twin model of the three zones in the goaf; Collect a number of historical multi-source monitoring data of the three zones in the goaf in the physical world corresponding to the digital twin model of the three zones in the digital world, and perform preprocessing to obtain a number of preprocessed historical multi-source monitoring data; According to the three-dimensional simulation model of the three zones in the goaf and a number of preprocessed historical multi-source monitoring data, train the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model of the digital twin model of the three zones in the goaf to obtain the final multi-source monitoring data analysis model and the final grouting dynamic regulation model.
3. The gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 2, characterized in that: The multi-source monitoring data analysis model is constructed based on the 3D-DBN-CNN-LSTM-ST-CNN algorithm, and the multi-source monitoring data analysis model includes a three-dimensional space feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm. The three-dimensional space feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module; The described grouting dynamic regulation model is constructed based on the MPO-MOGRPO-ICPO algorithm. The grouting dynamic regulation model includes a meta-strategy optimization module constructed based on the MPO algorithm, a regulation strategy generation module constructed based on the MOGRPO algorithm, and a regulation strategy optimization module constructed based on the ICPO algorithm, which are connected in sequence. The regulation strategy generation module includes an objective function set, an experience replay pool, an Actor network, and an agent. The agent is respectively connected to the objective function set, the experience replay pool, and the Actor network. The Actor network is connected to the regulation strategy optimization module.
4. The gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 3, characterized in that: Using deep learning algorithms, construct an initial multi-source monitoring data analysis model, and using reinforcement learning algorithms, construct an initial grouting dynamic regulation model, including the following steps: Use the 3D-DBN-CNN-LSTM-ST-CNN algorithm to construct an initial multi-source monitoring data analysis model; Use the MPO-MOGRPO-ICPO algorithm to construct an initial grouting dynamic regulation model; the initial grouting dynamic regulation model includes an initial meta-strategy optimization module, an initial regulation strategy generation module, and an initial regulation strategy optimization module; Set an objective function set, an experience replay pool, an Actor network, and an agent for the initial regulation strategy generation module, and use the initial network parameters of the Actor network as the output parameters of the initial meta-strategy optimization module; Take the grouting dynamic regulation strategy generation problem as the simulation environment of the initial regulation strategy generation module, and set the action space and state space for the agent of the initial regulation strategy generation module; Taking the minimization of the strategy generation error as the optimization goal, define the fitness function of the initial regulation strategy optimization module, and use the output of the Actor network of the initial regulation strategy generation module as the input of the initial regulation strategy optimization module.
5. A gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 4, characterized in that: According to the three-zone three-dimensional simulation model of the goaf and a number of preprocessed historical multi-source monitoring data, train the initial multi-source monitoring data analysis model and the initial grouting dynamic regulation model of the digital twin model of the three zones of the goaf to obtain the final multi-source monitoring data analysis model and the final grouting dynamic regulation model, including the following steps: Combine the three-zone three-dimensional simulation model of the goaf with a number of preprocessed historical multi-source monitoring data, and divide them into a model training set and a model test set according to a ratio of 7:3; According to the model training set, train the initial multi-source monitoring data analysis model of the digital twin model of the three zones of the goaf to obtain an optimized multi-source monitoring data analysis model, and generate a number of historical multi-source monitoring data analysis results; According to the model test set, test the optimized multi-source monitoring data analysis model. If the test accuracy is greater than the accuracy threshold, output the final multi-source monitoring data analysis model; otherwise, continue training; Traverse all the objective functions in the objective function set, and train the initial grouting dynamic regulation model according to a number of historical multi-source monitoring data analysis results to obtain the final grouting dynamic regulation model, and store the historical grouting dynamic regulation experience generated during the training process in the experience replay pool.
6. The gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 5, characterized in that: The historical multi-source monitoring data described above includes historical strata monitoring data, historical environmental monitoring data, and historical equipment operation monitoring data of the three zones in the goaf; The real-time multi-source monitoring data described above includes real-time strata monitoring data, real-time environmental monitoring data, and real-time equipment operation monitoring data of the three zones in the goaf.
7. A gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 6, characterized in that: Based on the digital twin model of the three zones in the goaf, collecting the real-time multi-source monitoring data of the three zones in the goaf and performing spatio-temporal correlation on the real-time multi-source monitoring data to obtain the real-time multi-source monitoring data after spatio-temporal correlation, including the following steps: Collect the real-time multi-source monitoring data of the three zones in the goaf in the physical world, perform preprocessing, and input the preprocessed real-time multi-source monitoring data into the digital twin model of the three zones in the goaf; Align the preprocessed real-time multi-source monitoring data in terms of time stamps to obtain the real-time multi-source monitoring data after time alignment; Perform spatial registration on the real-time multi-source monitoring data after time alignment to obtain the real-time multi-source monitoring data after spatial registration; Extract the real-time time features and real-time spatial features of the real-time multi-source monitoring data after spatial registration, and perform feature fusion on the real-time time features and real-time spatial features to obtain the real-time fused spatio-temporal features; Perform spatio-temporal correlation on the real-time fused spatio-temporal features in the digital twin model of the three zones in the goaf to obtain the real-time multi-source monitoring data after spatio-temporal correlation.
8. A gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 7, characterized in that: According to the real-time multi-source monitoring data after spatio-temporal correlation, using the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis result, including the following steps: Use the three-dimensional spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time three-dimensional spatial features of the three-dimensional simulation model of the three zones in the goaf in the digital twin model of the three zones in the goaf; Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the real-time multi-source monitoring data after spatio-temporal correlation; Use the structured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time structured data features of the real-time multi-source monitoring data after spatio-temporal correlation; According to the real-time three-dimensional spatial features, real-time unstructured data features, and real-time structured data features, use the multi-source monitoring data analysis module of the multi-source monitoring data analysis model to perform data analysis to obtain the real-time multi-source monitoring data analysis result.
9. A gob three-zone grouting dynamic regulation method based on reinforcement learning according to claim 8, characterized in that: According to the real-time multi-source monitoring data analysis result, using the grouting dynamic regulation model to generate a grouting dynamic regulation strategy to obtain the real-time grouting dynamic regulation strategy, including the following steps: According to the real-time multi-source monitoring data analysis result, use the meta-strategy optimization module of the grouting dynamic regulation model to adjust the Actor network of the regulation strategy generation module of the grouting dynamic regulation model to obtain the adjusted Actor network; Select the real-time objective function in the objective function set of the grouting dynamic regulation model, and randomly extract a number of historical grouting dynamic regulation experiences from the experience replay pool of the grouting dynamic regulation model based on the real-time objective function; According to the real-time multi-source monitoring data analysis result, adjust the state space of the intelligent agent to obtain the adjusted state space, and according to a number of historical grouting dynamic regulation experiences, adjust the action space of the intelligent agent to obtain the adjusted action space; Based on the real-time objective function, the adjusted state space, and the adjusted action space, an agent of the regulation strategy generation module is used to control the adjusted Actor network to generate the probability distribution of real-time grouting dynamic regulation actions; The regulation strategy optimization module of the grouting dynamic regulation model is used to optimize the probability distribution of real-time grouting dynamic regulation actions to obtain the real-time grouting dynamic regulation strategy.
10. A gob three-zone grouting dynamic regulation system based on reinforcement learning, which is used to implement the gob three-zone grouting dynamic regulation method as described in any one of claims 1-9, and is characterized in that: It includes a model construction unit, a spatio-temporal correlation unit, a data analysis unit, a regulation strategy generation unit, and a visualization unit connected in sequence.
Citation Information
Patent Citations
Coal mine comprehensive automatic management and control method and system based on digital twinborn technology
CN118462314A
Multi-source information advanced grouting reinforcement intelligent device and effect evaluation method
CN119777928A
Grain depot three-dimensional digital management system and method based on digital twinning
CN120031488A
Intelligent monitoring device for urban drainage system
CN120069791A
Substation equipment fault early warning method and system
CN120088959A
Cited By
Method and system for analyzing influence of goaf activation on tunnel stability
CN121859746A
An analysis method and system for the influence of goaf activation on tunnel stability
CN121859746B