A dynamic control method and system for three-zone grouting in goaf based on reinforcement learning

Through multi-source monitoring data analysis and digital twin technology based on reinforcement learning, the accuracy and applicability issues of the three-zone grouting technology in the goaf were solved, dynamic regulation of the three zones in the goaf was achieved, and the stability and applicability of the grouting effect were improved.

CN120258341BActive Publication Date: 2025-09-12POWERCHINA RAILWAY CONSTR +3
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510748346.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing three-zone grouting technology in goaf areas has problems such as insufficient accuracy, limited applicability, limited data processing capabilities and lack of dynamic optimization, resulting in unstable grouting effects.

Method used

A reinforcement learning-based method is used to construct a multi-source monitoring data analysis model and a grouting dynamic control model. Combined with digital twin technology, real-time multi-source monitoring data analysis and dynamic control of the three zones in the goaf are realized. Three-dimensional spatial features and unstructured data features are extracted through deep learning algorithms, and reinforcement learning algorithms are used to optimize grouting strategies and construct a grouting dynamic control system.

Benefits of technology

It improves the accuracy and efficiency of grouting control, adapts to complex geological conditions, provides a scientific basis for grouting material selection and strategy formulation, realizes comprehensive analysis of dynamic changes in time and space, and ensures the stability and reliability of grouting effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258341B_ABST
    Figure CN120258341B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of geotechnical engineering technology, and discloses a method and system for dynamic control of grouting in three zones of goaf based on reinforcement learning. The method comprises the following steps: constructing a multi-source monitoring data analysis model, a grouting dynamic control model and a digital twin model of three zones of goaf; based on the digital twin model of three zones of goaf, collecting real-time multi-source monitoring data and performing spatiotemporal correlation; using the multi-source monitoring data analysis model to perform data analysis based on the real-time multi-source monitoring data after spatiotemporal correlation; using the grouting dynamic control model to generate a grouting dynamic control strategy based on the analysis results of the real-time multi-source monitoring data; and using the digital twin model of three zones of goaf to visualize the real-time multi-source monitoring data analysis results and the real-time grouting dynamic control strategy. The present invention solves the problems of insufficient accuracy, limited applicability, limited data processing capabilities and lack of dynamic optimization in the existing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geotechnical engineering technology, and in particular relates to a method and system for dynamic control of three-zone grouting in goaf areas based on reinforcement learning. Background Art

[0002] Coal mine goafs refer to underground cavities or voids formed after coal resources are extracted during the mining process. The presence of such areas can significantly impact mine safety, the geological environment, and the lives of surrounding residents. Coal mine goafs are an inevitable occurrence during mining, and the geological and safety issues they present require significant attention. The three zones of goafs (collapse zone, fracture zone, and bend zone) are typical areas of rock movement formed during mining when the overlying rock strata lose support and deform and fail due to the extraction of underground coal. Grouting in the three zones of goafs involves using grouting equipment to inject slurry (such as cement slurry or coal gangue slurry) into the collapse zone, fracture zone, and bend zone to fill the cavities, strengthen the rock strata, and improve geological conditions. This method can effectively address problems such as ground collapse and mine water inrush caused by goafs, and is suitable for reinforcement in complex strata such as mines and tunnels.

[0003] In the field of three-zone grouting technology in goaf, although some technologies and methods have been applied, they still have many defects, as follows:

[0004] 1) Insufficient precision: Traditional grouting technology usually relies on empirical judgment, making it difficult to achieve precise control of the grouting range. For example, when filling cracks in goaf areas, inaccurate grouting ranges may result in some areas not being effectively filled, or slurry may leak into non-target areas, resulting in resource waste.

[0005] 2) Limited applicability: For complex geological conditions (such as high stress and high permeability areas), traditional grouting methods are difficult to meet the requirements of efficient filling, and the selection of grouting materials lacks specificity;

[0006] 3) Limited data processing capabilities: Existing monitoring data are mostly static analysis, lacking the ability to comprehensively analyze spatiotemporal dynamic changes, and unable to provide real-time decision support for dynamic grouting control;

[0007] 4) Lack of dynamic optimization: Existing grouting control strategies are mostly static planning, which cannot dynamically adjust grouting decisions (such as grouting volume, grouting location, and grouting pressure) based on real-time monitoring data, resulting in unstable grouting effects. Summary of the Invention

[0008] In order to solve the problems of insufficient accuracy, limited applicability, limited data processing capabilities and lack of dynamic optimization in the existing technology, the purpose of the present invention is to provide a dynamic control method and system for three-zone grouting in goaf based on reinforcement learning.

[0009] The technical solution adopted in the present invention is:

[0010] A method for dynamic control of three-zone grouting in goaf based on reinforcement learning, comprising the following steps:

[0011] Using deep learning algorithms, a multi-source monitoring data analysis model was constructed. Using reinforcement learning algorithms, a grouting dynamic control model was constructed. And using digital twin technology, a three-zone digital twin model of the goaf was constructed.

[0012] Based on the real-time multi-source monitoring data after temporal and spatial correlation, a multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results;

[0013] Based on the analysis results of real-time multi-source monitoring data, the grouting dynamic control model is used to generate a grouting dynamic control strategy to obtain a real-time grouting dynamic control strategy;

[0014] Using the three-zone digital twin model of the goaf, the real-time multi-source monitoring data analysis results and the real-time grouting dynamic control strategy are visualized, the real-time grouting dynamic control strategy is executed, and the data collection step continues.

[0015] Furthermore, a deep learning algorithm was used to build a multi-source monitoring data analysis model, a reinforcement learning algorithm was used to build a grouting dynamic control model, and digital twin technology was used to build a three-zone digital twin model of the goaf, including the following steps:

[0016] Use deep learning algorithms to build an initial multi-source monitoring data analysis model, and use reinforcement learning algorithms to build an initial grouting dynamic control model;

[0017] Construct a three-dimensional simulation model of the three zones in the goaf. Combined with the initial multi-source monitoring data analysis model and the initial grouting dynamic control model, digital twin technology is used to construct a digital twin model of the three zones in the goaf.

[0018] Collecting a number of historical multi-source monitoring data of the three zones of the goaf in the physical world corresponding to the digital twin model of the three zones of the goaf in the digital world, and preprocessing them to obtain a number of preprocessed historical multi-source monitoring data;

[0019] Based on the three-dimensional simulation model of the three zones in the goaf and several pre-processed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic control model of the digital twin model of the three zones in the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic control model.

[0020] Furthermore, the multi-source monitoring data analysis model is constructed based on the 3D-DBN-CNN-LSTM-ST-CNN algorithm, and the multi-source monitoring data analysis model includes a three-dimensional spatial feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm, and the three-dimensional spatial feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module;

[0021] The grouting dynamic control model is constructed based on the MPO-MOGRPO-ICPO algorithm, and the grouting dynamic control model includes a meta-strategy optimization module constructed based on the MPO algorithm, a control strategy generation module constructed based on the MOGRPO algorithm, and a control strategy optimization module constructed based on the ICPO algorithm, which are connected in sequence. The control strategy generation module includes an objective function set, an experience replay pool, an Actor network and an intelligent agent, and the intelligent agent is connected to the objective function set, the experience replay pool and the Actor network respectively, and the Actor network is connected to the control strategy optimization module.

[0022] Furthermore, a deep learning algorithm is used to construct an initial multi-source monitoring data analysis model, and a reinforcement learning algorithm is used to construct an initial grouting dynamic control model, including the following steps:

[0023] Use the 3D-DBN-CNN-LSTM-ST-CNN algorithm to build an initial multi-source monitoring data analysis model;

[0024] The MPO-MOGRPO-ICPO algorithm is used to construct an initial grouting dynamic control model; the initial grouting dynamic control model includes an initial meta-strategy optimization module, an initial control strategy generation module, and an initial control strategy optimization module;

[0025] Set the objective function set, experience replay pool, actor network, and agent for the initial control strategy generation module, and use the initial network parameters of the actor network as the output parameters of the initial meta-strategy optimization module;

[0026] The grouting dynamic control strategy generation problem is used as the simulation environment of the initial control strategy generation module, and the action space and state space are set for the intelligent agent of the initial control strategy generation module;

[0027] Taking minimizing the strategy generation error as the optimization goal, the fitness function of the initial control strategy optimization module is defined, and the output of the Actor network of the initial control strategy generation module is used as the input of the initial control strategy optimization module.

[0028] Furthermore, based on the three-dimensional simulation model of the three zones in the goaf and some pre-processed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic control model of the three-zone digital twin model of the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic control model, including the following steps:

[0029] The three-dimensional simulation model of the goaf is combined with some pre-processed historical multi-source monitoring data, and divided into model training set and model test set in a ratio of 7:3.

[0030] Based on the model training set, the initial multi-source monitoring data analysis model of the three-zone digital twin model of the goaf is trained to obtain an optimized multi-source monitoring data analysis model and generate several historical multi-source monitoring data analysis results;

[0031] The optimized multi-source monitoring data analysis model is tested based on the model test set. If the test accuracy is greater than the accuracy threshold, the final multi-source monitoring data analysis model is output; otherwise, training continues.

[0032] Traverse all the objective functions in the objective function set, train the initial grouting dynamic control model based on the analysis results of several historical multi-source monitoring data, obtain the final grouting dynamic control model, and store the historical grouting dynamic control experience generated during the training process in the experience replay pool.

[0033] Furthermore, the historical multi-source monitoring data includes historical stratigraphic monitoring data of the three zones of the goaf, historical environmental monitoring data, and historical equipment operation monitoring data;

[0034] Real-time multi-source monitoring data includes real-time stratum monitoring data of three zones in the goaf, real-time environmental monitoring data, and real-time equipment operation monitoring data.

[0035] Furthermore, based on the digital twin model of the three zones in the goaf, real-time multi-source monitoring data of the three zones in the goaf are collected, and the real-time multi-source monitoring data are temporally and spatially correlated to obtain the real-time multi-source monitoring data after temporal and spatial correlation, including the following steps:

[0036] Collect real-time multi-source monitoring data of the three zones of the goaf in the physical world, perform preprocessing, and input the obtained preprocessed real-time multi-source monitoring data into the digital twin model of the three zones of the goaf;

[0037] Align the pre-processed real-time multi-source monitoring data on the timestamp to obtain the time-aligned real-time multi-source monitoring data;

[0038] Perform spatial registration on the time-aligned real-time multi-source monitoring data to obtain spatially registered real-time multi-source monitoring data;

[0039] Extract the real-time time features and real-time spatial features of the real-time multi-source monitoring data after spatial registration, and perform feature fusion on the real-time time features and real-time spatial features to obtain real-time fused spatiotemporal features;

[0040] The real-time fusion of spatiotemporal features is spatiotemporally correlated in the three-zone digital twin model of the goaf to obtain real-time multi-source monitoring data after spatiotemporal correlation.

[0041] Furthermore, based on the real-time multi-source monitoring data after temporal and spatial correlation, a multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results, including the following steps:

[0042] Use the 3D spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time 3D spatial features of the 3D simulation model of the goaf in the 3D digital twin model of the goaf;

[0043] Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the real-time multi-source monitoring data after temporal and spatial correlation;

[0044] Use the structured data feature extraction module of the multi-source monitoring data analysis model to extract real-time structured data features of real-time multi-source monitoring data after temporal and spatial correlation;

[0045] According to the real-time three-dimensional spatial characteristics, real-time unstructured data characteristics and real-time structured data characteristics, the multi-source monitoring data analysis module of the multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results.

[0046] Furthermore, based on the analysis results of the real-time multi-source monitoring data, a grouting dynamic control model is used to generate a grouting dynamic control strategy, and a real-time grouting dynamic control strategy is obtained, which includes the following steps:

[0047] According to the analysis results of real-time multi-source monitoring data, the meta-strategy optimization module of the grouting dynamic control model is used to adjust the Actor network of the control strategy generation module of the grouting dynamic control model to obtain the adjusted Actor network;

[0048] A real-time objective function is selected from the objective function set of the grouting dynamic control model, and based on the real-time objective function, a number of historical grouting dynamic control experiences are randomly sampled from the experience playback pool of the grouting dynamic control model;

[0049] According to the analysis results of real-time multi-source monitoring data, the state space of the intelligent agent is adjusted to obtain the adjusted state space. According to some historical grouting dynamic control experience, the action space of the intelligent agent is adjusted to obtain the adjusted action space.

[0050] Based on the real-time objective function, the adjusted state space, and the adjusted action space, the intelligent agent of the control strategy generation module is used to control the adjusted Actor network to generate the real-time grouting dynamic control action probability distribution;

[0051] The control strategy optimization module of the grouting dynamic control model is used to optimize the probability distribution of real-time grouting dynamic control actions to obtain the real-time grouting dynamic control strategy.

[0052] A reinforcement learning-based dynamic control system for three-zone grouting in goafs is used to implement a dynamic control method for three-zone grouting in goafs, comprising a model construction unit, a time-space association unit, a data analysis unit, a control strategy generation unit, and a visualization unit connected in sequence.

[0053] The beneficial effects of the present invention are:

[0054] The present invention provides a method and system for dynamic grouting control of three zones in goaf based on reinforcement learning. The multi-source monitoring data analysis model constructed avoids reliance on empirical judgment, comprehensively considers the information on strata, environment and equipment operation included in the multi-source monitoring data, and combines three-dimensional spatial information such as the geometric shape, spatial distribution and volume change of the three zones in the goaf, thereby improving the accuracy and efficiency of grouting control and avoiding waste of resources; it can comprehensively analyze multi-source monitoring data under complex geological conditions, provide a scientific basis for the selection of grouting materials and the formulation of grouting strategies, and at the same time, can dynamically adjust grouting parameters according to different geological conditions to improve applicability in different geological environments; it combines deep learning algorithms and reinforcement learning algorithms to realize automated multi-source monitoring data analysis, improve data processing efficiency and accuracy, and has the ability to comprehensively analyze dynamic changes in time and space, providing real-time decision support for dynamic grouting control; the constructed dynamic grouting control model adopts a dynamic optimization mechanism, which can timely adjust grouting decisions according to the real-time multi-source monitoring data analysis results, thereby improving the stability and reliability of the grouting effect.

[0055] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flow chart of the dynamic control method of three-zone grouting in goaf based on reinforcement learning in the present invention.

[0057] Figure 2It is a structural block diagram of the three-zone grouting dynamic control system of the goaf based on reinforcement learning in the present invention. DETAILED DESCRIPTION

[0058] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0059] Example 1:

[0060] like Figure 1 As shown, this embodiment provides a method for dynamic control of three-zone grouting in goaf based on reinforcement learning, comprising the following steps:

[0061] S1: Use deep learning algorithms to build a multi-source monitoring data analysis model, use reinforcement learning algorithms to build a grouting dynamic control model, and use digital twin technology to build a three-zone digital twin model of the goaf, including the following steps:

[0062] S1-1: Use deep learning algorithms to build an initial multi-source monitoring data analysis model, and use reinforcement learning algorithms to build an initial grouting dynamic control model;

[0063] The multi-source monitoring data analysis model is constructed based on the 3D-Deep Belief Network (DBN)-Convolutional Neural Network (CNN)-Long Short-Term Memory Network (LSTM)-Spatio-Temporal Convolutional Neural Network (ST-CNN) algorithm. The multi-source monitoring data analysis model includes a three-dimensional spatial feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm. The three-dimensional spatial feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module.

[0064] The 3D-DBN of the three-dimensional spatial feature extraction module is a deep learning algorithm that can extract complex spatial features from three-dimensional simulation models (such as three-dimensional data, three-dimensional grids or voxel data), improve the extraction accuracy of three-dimensional spatial features, and provide high-quality three-dimensional features for subsequent multi-source data analysis. Through deep learning technology, it can process complex spatial distribution and structural information and enhance the representation ability of the model; the CNN of the unstructured data feature extraction module is good at processing unstructured data such as images and videos, and can extract local features and hierarchical information therein, improve the analysis ability of unstructured data (such as images, videos, etc.), and effectively capture the key information therein. Through the hierarchical feature extraction of CNN, it can better identify patterns in the data, which is Provide support for subsequent analysis; the LSTM of the structured data feature extraction module is a variant of the recurrent neural network, which is good at processing time series data and can capture long-term dependencies in the data, improving the processing ability of time series data, especially when the data has long-term dependencies. Through the time series feature extraction of LSTM, it can better understand the dynamic changes of the data and provide support for data analysis; the ST-CNN of the multi-source monitoring data analysis module is a deep learning algorithm that combines spatial and temporal information. It can process spatiotemporal data and extract its features. Through the fusion of spatiotemporal features, it can comprehensively analyze the dynamic changes of the three zones in the goaf, improve the accuracy and real-time performance of data analysis, and provide reliable data support for tasks such as dynamic grouting control;

[0065] The grouting dynamic control model is constructed based on the Meta-Policy Optimization (MPO)-Multi-Objective Group Relative Policy Optimization (MOGRPO)-Improved Crested Porcupine Optimizer (ICPO) algorithm. The grouting dynamic control model includes a meta-policy optimization module based on the MPO algorithm, a control strategy generation module based on the MOGRPO algorithm, and a control strategy optimization module based on the ICPO algorithm, which are connected in sequence. The control strategy generation module includes an objective function set, an experience replay pool, an actor network, and an intelligent agent. The intelligent agent is connected to the objective function set, the experience replay pool, and the actor network respectively, and the actor network is connected to the control strategy optimization module.

[0066] The meta-strategy optimization module is used to control the initial network parameters of the Actor network in the strategy generation module so that these parameters can quickly adapt to new, unseen multi-source monitoring data analysis results, thereby improving the generalization ability of the model. Even under unseen multi-source monitoring data analysis results, the Actor network can be updated based on previous learning experience, thereby improving the adaptability of the control strategy generation model. The objective function set of the control strategy generation module can handle multiple conflicting optimization goals, such as equipment operating efficiency, parameter control cost, operating safety, etc., and generate control strategies that balance these goals. The intelligent agent learns historical experience through the experience replay pool and continuously optimizes its own strategy generation ability. The intelligent agent controls the Actor network based on the learned experience to generate a more effective grouting dynamic control strategy. The design of the experience replay pool and the intelligent agent enables the model to continuously learn and optimize, thereby improving the quality of strategy generation. Since the control strategy generation module adopts a group exploration method, it can avoid falling into local optimal solutions to a certain extent. , the Actor network outputs the distribution probability of actions in a given state. The goal is to learn an optimal strategy, that is, to maximize the long-term cumulative reward. In the continuous action space, the Actor network usually outputs a mean and an optional variance parameter to describe the probability distribution of the action. The Critic network is responsible for evaluating the value of a given state, that is, predicting the expected return that can be obtained by starting from this state and following the current strategy. It usually outputs a scalar value to represent the value of the state or the state-action value. The experience replay pool is used to store historical experience for reuse during training. The control strategy generation module directly updates the Actor network through gradients, eliminating the Critic network in traditional reinforcement learning, making the algorithm structure simpler. The ICPO algorithm of the control strategy optimization module is an efficient optimization algorithm that can quickly find the optimal or approximately optimal solution. By optimizing the distribution probability of actions in a given state output by the Actor network, the accuracy and generation efficiency of the real-time grouting dynamic control strategy are further improved.

[0067] Using deep learning algorithms, we build an initial multi-source monitoring data analysis model. Using reinforcement learning algorithms, we build an initial grouting dynamic control model. This includes the following steps:

[0068] S1-1-1: Use the 3D-DBN-CNN-LSTM-ST-CNN algorithm to build an initial multi-source monitoring data analysis model;

[0069] S1-1-2: Use the MPO-MOGRPO-ICPO algorithm to construct an initial grouting dynamic control model; the initial grouting dynamic control model includes an initial meta-strategy optimization module, an initial control strategy generation module, and an initial control strategy optimization module;

[0070] S1-1-3: Set the objective function set, experience replay pool, actor network, and agent for the initial control strategy generation module, and use the initial network parameters of the actor network as the output parameters of the initial meta-strategy optimization module;

[0071] S1-1-4: Use the grouting dynamic control strategy generation problem as the simulation environment for the initial control strategy generation module, and set the action space and state space for the intelligent agent of the initial control strategy generation module;

[0072] S1-1-5: With minimizing the strategy generation error as the optimization goal, define the fitness function of the initial control strategy optimization module, and use the output of the Actor network of the initial control strategy generation module as the input of the initial control strategy optimization module;

[0073] S1-2: Construct a three-dimensional simulation model of the three zones in the goaf. Combined with the initial multi-source monitoring data analysis model and the initial grouting dynamic control model, use digital twin technology to construct a digital twin model of the three zones in the goaf. The model includes the following steps:

[0074] S1-2-1: Collect geological data, Geographic Information System (GIS) data, and 3D scanning data of the three zones of the goaf. Based on the GIS data and 3D scanning data, use 3D simulation technology to construct a 3D simulation model of the three zones of the goaf.

[0075] S1-2-2: Use digital twin technology to digitally map the three-dimensional simulation model of the goaf’s three zones to obtain a digital twin framework of the goaf’s three zones;

[0076] S1-2-3: Set the initial multi-source monitoring data analysis model and the initial grouting dynamic control model in the three-zone digital twin framework of the goaf;

[0077] S1-2-4: Convert the real-time multi-source monitoring data into a real-time data stream, and connect the real-time data stream to the initial multi-source monitoring data analysis model and the initial grouting dynamic control model in the three-zone digital twin framework of the goaf to obtain a three-dimensional simulation model of the three-zone goaf;

[0078] S1-3: Collect a number of historical multi-source monitoring data of the three zones of the goaf in the physical world corresponding to the digital twin model of the three zones of the goaf in the digital world, and pre-process them to obtain a number of pre-processed historical multi-source monitoring data;

[0079] Historical multi-source monitoring data includes historical stratigraphic monitoring data of three zones in the goaf, historical environmental monitoring data, and historical equipment operation monitoring data;

[0080] The three zones of the goaf include the collapse zone, the fracture zone and the bending zone;

[0081] The stratum monitoring data include: the collapse height of the rock formation in the collapse zone (the maximum height of the collapse zone and its change over time), the distribution and particle size of the collapsed rock blocks (monitoring the distribution and particle size of the collapsed rock blocks through borehole photography or geological radar), the rock formation movement speed (using monitoring equipment to record the speed and dynamic changes of the rock formation during the collapse process), and the collapse range (the horizontal extension of the collapse zone);

[0082] Fracture development height of the fracture zone (maximum fracture zone height and its change over time), fracture width and density (fracture width and density detected through borehole photography or ultrasonic wave detection), fracture direction (fracture direction and inclination, used to assess rock formation stability), water seepage (water seepage volume and flow direction within the fracture zone, used to determine whether groundwater contamination may occur);

[0083] Bending deformation of the bending zone (monitoring the bending degree of the rock formation through strain gauges or inclinometers), bending zone thickness (the thickness of the bending zone from the top of the fracture zone to the unaffected rock formation), rock formation displacement (horizontal and vertical displacement of the rock formation within the bending zone), stress distribution (monitoring the stress changes within the bending zone through strain gauges);

[0084] Environmental monitoring data include the ambient temperature, humidity, altitude, longitude and latitude of collapse zones, fracture zones and bending zones;

[0085] Equipment operation monitoring data includes the operating power, operating current, operating voltage, operating mode, operating parameters, and fault codes of the grouting equipment;

[0086] Preprocessing includes data cleaning, denoising, and normalization to ensure data quality and provide data support for subsequent model training and optimization;

[0087] S1-4: Based on the three-dimensional simulation model of the three zones in the goaf and some pre-processed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic control model of the three-zone digital twin model of the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic control model, including the following steps:

[0088] S1-4-1: Combine the three-dimensional simulation model of the goaf with several pre-processed historical multi-source monitoring data, and divide them into model training set and model test set in a ratio of 7:3;

[0089] S1-4-2: Based on the model training set, the initial multi-source monitoring data analysis model of the three-zone digital twin model of the goaf is trained to obtain an optimized multi-source monitoring data analysis model and generate several historical multi-source monitoring data analysis results;

[0090] S1-4-3: Test the optimized multi-source monitoring data analysis model based on the model test set. If the test accuracy is greater than the accuracy threshold, output the final multi-source monitoring data analysis model; otherwise, continue training.

[0091] S1-4-4: Traverse all objective functions in the objective function set, train the initial grouting dynamic control model based on the analysis results of several historical multi-source monitoring data, obtain the final grouting dynamic control model, and store the historical grouting dynamic control experience generated during the training process in the experience replay pool;

[0092] S2: Based on the digital twin model of the three zones in the goaf, real-time multi-source monitoring data of the three zones in the goaf is collected, and the real-time multi-source monitoring data is temporally and spatially correlated to obtain the real-time multi-source monitoring data after temporal and spatial correlation, including the following steps:

[0093] S2-1: Collect real-time multi-source monitoring data of the three zones of the goaf in the physical world, perform preprocessing, and input the obtained preprocessed real-time multi-source monitoring data into the digital twin model of the three zones of the goaf;

[0094] Real-time multi-source monitoring data includes real-time stratum monitoring data of three zones in the goaf, real-time environmental monitoring data, and real-time equipment operation monitoring data;

[0095] S2-2: Align the pre-processed real-time multi-source monitoring data based on the timestamp to obtain the time-aligned real-time multi-source monitoring data;

[0096] S2-3: spatially register the time-aligned real-time multi-source monitoring data to obtain spatially registered real-time multi-source monitoring data;

[0097] S2-4: extracting the real-time temporal features and real-time spatial features of the real-time multi-source monitoring data after spatial registration, and performing feature fusion on the real-time temporal features and real-time spatial features to obtain real-time fused spatiotemporal features;

[0098] S2-5: Perform spatiotemporal correlation on the real-time fused spatiotemporal features in the three-zone digital twin model of the goaf to obtain real-time multi-source monitoring data after spatiotemporal correlation;

[0099] S3: Based on the real-time multi-source monitoring data after temporal and spatial correlation, a multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results, including the following steps:

[0100] S3-1: Use the 3D spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time 3D spatial features of the 3D simulation model of the goaf three zones in the goaf three zones digital twin model;

[0101] S3-2: Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the real-time multi-source monitoring data after temporal and spatial correlation;

[0102] S3-3: Use the structured data feature extraction module of the multi-source monitoring data analysis model to extract real-time structured data features of real-time multi-source monitoring data after temporal and spatial correlation;

[0103] S3-4: Based on the real-time three-dimensional spatial features, the real-time unstructured data features, and the real-time structured data features, a multi-source monitoring data analysis module of the multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results;

[0104] Real-time multi-source monitoring data analysis results include real-time grouting equipment abnormality analysis results, real-time grouting demand analysis results, real-time grouting position analysis results, etc.

[0105] S4: Based on the analysis results of the real-time multi-source monitoring data, a grouting dynamic control model is used to generate a grouting dynamic control strategy to obtain a real-time grouting dynamic control strategy, including the following steps:

[0106] S4-1: Based on the analysis results of real-time multi-source monitoring data, the meta-strategy optimization module of the grouting dynamic control model is used to adjust the Actor network of the control strategy generation module of the grouting dynamic control model to obtain the adjusted Actor network;

[0107] S4-2: Select a real-time objective function from the objective function set of the grouting dynamic control model, and based on the real-time objective function, randomly extract a number of historical grouting dynamic control experiences from the experience playback pool of the grouting dynamic control model;

[0108] S4-3: Based on the analysis results of real-time multi-source monitoring data, the state space of the intelligent agent is adjusted to obtain the adjusted state space. Based on some historical grouting dynamic control experience, the action space of the intelligent agent is adjusted to obtain the adjusted action space.

[0109] S4-4: Based on the real-time objective function, the adjusted state space, and the adjusted action space, the intelligent agent of the control strategy generation module is used to control the adjusted Actor network to generate the real-time grouting dynamic control action probability distribution, including the following steps:

[0110] S4-4-1: Based on the real-time objective function, the agent of the control strategy generation module is used to control the adjusted actor network to generate the probability distribution of all possible actions in the adjusted action space corresponding to each real-time state in the adjusted state space;

[0111] S4-4-2: Integrate the probability distribution of all possible actions in the same real-time state to obtain a real-time probability distribution sequence;

[0112] S4-4-2: Traverse all real-time states in the adjusted state space to obtain a real-time grouting dynamic control action probability distribution consisting of several real-time probability distribution sequences;

[0113] S4-5: Using the control strategy optimization module of the grouting dynamic control model, the probability distribution of the real-time grouting dynamic control action is optimized to obtain the real-time grouting dynamic control strategy, including the following steps:

[0114] S4-5-1: Encode the real-time optimization parameter values ​​of the real-time grouting dynamic control action probability distribution into the solution vector of the control strategy optimization module, and set the ICPO population parameters and the maximum number of iterations;

[0115] S4-5-2: Initialize according to the solution vector and ICPO population parameters to obtain several initial solutions; the initial solutions correspond to the initial real-time optimization parameter values;

[0116] The formula is:

[0117]

[0118] Where, is the initial ICPO individual of the Circle chaos map, i.e. the initial solution; is the randomly generated initial ICPO individual; is the ICPO individual indicator; compared with the randomly distributed population, the initial position distribution of the improved ICPO individuals generated by the Circle Chaotic Map sequence is more uniform, which expands the search range of the algorithm in space, increases the diversity of group positions, and to a certain extent improves the defect of the algorithm being prone to falling into local extreme values, thereby improving the optimization efficiency of the algorithm;

[0119] S4-5-3: Introduce a cyclic population reduction mechanism to limit the number of individuals in the ICPO population parameters and obtain the updated ICPO population parameters for the next iteration;

[0120] The formula is:

[0121]

[0122] Where, For the The number of individuals in the ICPO population parameter of the iteration; For the The number of individuals in the ICPO population parameter of the iteration; is the minimum number of individuals in the ICPO population parameter; Evaluate arguments for functions; Evaluate loop parameters for a function; Evaluate loop parameters for the maximum function; t is the indicator of the number of iterations;

[0123] S4-5-4: Calculate the initial fitness value of the initial ICPO individuals in the initial ICPO population according to the fitness function;

[0124] The formula for fitness value is:

[0125]

[0126] Where, is the fitness function; is the mean square error function, which is used to obtain the strategy generation error; For the The true output value of the sample; For the The ideal output value of samples; is the sample indicator amount; N is the total number of samples;

[0127] S4-5-5: Based on the initial fitness value and the updated ICPO population parameters, the first defense strategy, the second defense strategy, the third defense strategy, and the fourth defense strategy are used to update the initial ICPO population to obtain an updated ICPO population;

[0128] The formula for the first defense strategy is:

[0129]

[0130] Where, For the updated ICPO individuals within the first defense range; The initial ICPO individual within the first defense range; is a random number based on normal distribution; is a random value in the interval [0,1]; It is the optimal solution within the first defense range; is the vector generated between the true optimal solution within the first defense range and the optimal solution randomly selected from the ICPO population; It is the individual indicator of ICPO; is the iteration indicator;

[0131] The formula for the second defense strategy is:

[0132]

[0133] Where, For the updated ICPO individuals within the second defense range; The initial ICPO individual within the second defense range; is the search upper limit vector of the second defense range; is a random value in the interval [0,1]; Respectively Initial ICPO individuals; All are [1, ] two random integers between; is the vector generated between the true optimal solution within the second defense range and the optimal solution randomly selected from the ICPO population;

[0134] The formula for the third defense strategy is:

[0135]

[0136] Where, For the updated ICPO individuals within the third defense range; The initial ICPO individual within the third defense range; is the search upper limit vector of the third defense range; Respectively Initial ICPO individuals; is [1, ] a random integer between ; The odor diffusion factor defined for the fitness function; It is a defense factor; Control parameters for search direction;

[0137] The formula for the fourth defense strategy is:

[0138]

[0139] Where, For the updated ICPO individuals within the fourth defense range; For the initial ICPO individuals within the fourth defense range; It is the optimal solution within the fourth defense range; All are random values ​​in the interval [0,1]; It is a defense factor; Control parameters for search direction; is the average force affecting the search direction; is the convergence speed factor;

[0140] S4-5-6: Use the dynamic reverse learning algorithm to perform dynamic reverse learning on the updated ICPO population to generate a dynamic reverse ICPO population;

[0141] The formula is:

[0142]

[0143] Where, It is a dynamically reversed ICPO individual; is the decreasing inertia coefficient; are the maximum and minimum values ​​of the vector space respectively; For the updated ICPO individual;

[0144] S4-5-7: According to the fitness function, calculate the fitness values ​​of all ICPO individuals in the updated ICPO population and the dynamically reversed ICPO population, take the ICPO individual with the minimum fitness value as the optimal individual, and retain the optimal individual;

[0145] S4-5-8: If the number of iterations of iterative optimization reaches the maximum number of iterations or the fitness value of the optimal individual meets the requirements, the optimal solution corresponding to the optimal individual is output, and the solution vector of the optimal solution is decoded to obtain the optimal real-time optimization parameter value;

[0146] S4-5-9: Based on the optimal real-time optimization parameter values, the probability distribution of the real-time grouting dynamic control action is optimized to obtain the real-time grouting dynamic control strategy;

[0147] Real-time grouting dynamic control strategies include real-time operation parameter adjustment decisions of grouting equipment, real-time grouting pressure control decisions, real-time grouting type control decisions, real-time grouting volume control decisions, etc.

[0148] S5: Use the three-zone digital twin model of the goaf to visualize the real-time multi-source monitoring data analysis results and the real-time grouting dynamic control strategy, execute the real-time grouting dynamic control strategy, and continue the data collection step.

[0149] Example 2:

[0150] like Figure 2 As shown, this embodiment provides a three-zone grouting dynamic control system for goaf areas based on reinforcement learning, which is used to implement a three-zone grouting dynamic control method for goaf areas, including a model construction unit, a time-space correlation unit, a data analysis unit, a control strategy generation unit and a visualization unit connected in sequence.

[0151] The model building unit is used to build a multi-source monitoring data analysis model using a deep learning algorithm, a grouting dynamic control model using a reinforcement learning algorithm, and a digital twin model of the three zones of the goaf using digital twin technology;

[0152] The spatiotemporal correlation unit is used to collect real-time multi-source monitoring data of the three zones of the goaf based on the digital twin model of the three zones of the goaf, and to perform spatiotemporal correlation on the real-time multi-source monitoring data to obtain the real-time multi-source monitoring data after spatiotemporal correlation;

[0153] A data analysis unit is used to perform data analysis based on the real-time multi-source monitoring data after temporal and spatial correlation using a multi-source monitoring data analysis model to obtain real-time multi-source monitoring data analysis results;

[0154] A control strategy generation unit is used to generate a grouting dynamic control strategy based on the analysis results of real-time multi-source monitoring data using a grouting dynamic control model to obtain a real-time grouting dynamic control strategy;

[0155] The visualization unit is used to use the three-zone digital twin model of the goaf to visualize the real-time multi-source monitoring data analysis results and the real-time grouting dynamic control strategy, execute the real-time grouting dynamic control strategy, and continue the data collection step.

[0156] The present invention provides a method and system for dynamic grouting control of three zones in goaf based on reinforcement learning. The multi-source monitoring data analysis model constructed avoids reliance on empirical judgment, comprehensively considers the information on strata, environment and equipment operation included in the multi-source monitoring data, and combines three-dimensional spatial information such as the geometric shape, spatial distribution and volume change of the three zones in the goaf, thereby improving the accuracy and efficiency of grouting control and avoiding waste of resources; it can comprehensively analyze multi-source monitoring data under complex geological conditions, provide a scientific basis for the selection of grouting materials and the formulation of grouting strategies, and at the same time, can dynamically adjust grouting parameters according to different geological conditions to improve applicability in different geological environments; it combines deep learning algorithms and reinforcement learning algorithms to realize automated multi-source monitoring data analysis, improve data processing efficiency and accuracy, and has the ability to comprehensively analyze dynamic changes in time and space, providing real-time decision support for dynamic grouting control; the constructed dynamic grouting control model adopts a dynamic optimization mechanism, which can timely adjust grouting decisions according to the real-time multi-source monitoring data analysis results, thereby improving the stability and reliability of the grouting effect.

[0157] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.

Claims

1. A dynamic control method for three-zone grouting in goaf based on reinforcement learning, characterized by: The steps include: Using deep learning algorithms, a multi-source monitoring data analysis model was constructed. Using reinforcement learning algorithms, a grouting dynamic control model was constructed. And using digital twin technology, a three-zone digital twin model of the goaf was constructed. The multi-source monitoring data analysis model is constructed based on the 3D-DBN-CNN-LSTM-ST-CNN algorithm, and the multi-source monitoring data analysis model includes a three-dimensional spatial feature extraction module constructed based on the 3D-DBN algorithm, an unstructured data feature extraction module constructed based on the CNN algorithm, a structured data feature extraction module constructed based on the LSTM, and a multi-source monitoring data analysis module constructed based on the ST-CNN algorithm. The three-dimensional spatial feature extraction module, the unstructured data feature extraction module, and the structured data feature extraction module are all connected to the multi-source monitoring data analysis module; The grouting dynamic control model is constructed based on the MPO-MOGRPO-ICPO algorithm, and the grouting dynamic control model includes a meta-strategy optimization module constructed based on the MPO algorithm, a control strategy generation module constructed based on the MOGRPO algorithm, and a control strategy optimization module constructed based on the ICPO algorithm, which are connected in sequence. The control strategy generation module includes an objective function set, an experience replay pool, an Actor network, and an intelligent agent, and the intelligent agent is respectively connected to the objective function set, the experience replay pool, and the Actor network, and the Actor network is connected to the control strategy optimization module; Based on the digital twin model of the three zones in the goaf, real-time multi-source monitoring data of the three zones in the goaf is collected, and the real-time multi-source monitoring data is temporally and spatially correlated to obtain the real-time multi-source monitoring data after temporal and spatial correlation; Based on the real-time multi-source monitoring data after temporal and spatial correlation, a multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results, including the following steps: Use the 3D spatial feature extraction module of the multi-source monitoring data analysis model to extract the real-time 3D spatial features of the 3D simulation model of the goaf in the 3D digital twin model of the goaf; Use the unstructured data feature extraction module of the multi-source monitoring data analysis model to extract the real-time unstructured data features of the real-time multi-source monitoring data after temporal and spatial correlation; Use the structured data feature extraction module of the multi-source monitoring data analysis model to extract real-time structured data features of real-time multi-source monitoring data after temporal and spatial correlation; According to the real-time three-dimensional spatial features, real-time unstructured data features and real-time structured data features, the multi-source monitoring data analysis module of the multi-source monitoring data analysis model is used to perform data analysis to obtain real-time multi-source monitoring data analysis results; Based on the analysis results of real-time multi-source monitoring data, the grouting dynamic control model is used to generate a grouting dynamic control strategy, and the real-time grouting dynamic control strategy is obtained, which includes the following steps: According to the analysis results of real-time multi-source monitoring data, the meta-strategy optimization module of the grouting dynamic control model is used to adjust the Actor network of the control strategy generation module of the grouting dynamic control model to obtain the adjusted Actor network; A real-time objective function is selected from the objective function set of the grouting dynamic control model, and based on the real-time objective function, a number of historical grouting dynamic control experiences are randomly sampled from the experience playback pool of the grouting dynamic control model; According to the analysis results of real-time multi-source monitoring data, the state space of the intelligent agent is adjusted to obtain the adjusted state space. According to some historical grouting dynamic control experience, the action space of the intelligent agent is adjusted to obtain the adjusted action space. Based on the real-time objective function, the adjusted state space, and the adjusted action space, the intelligent agent of the control strategy generation module is used to control the adjusted Actor network to generate the real-time grouting dynamic control action probability distribution; Using the control strategy optimization module of the grouting dynamic control model, the probability distribution of real-time grouting dynamic control actions is optimized to obtain the real-time grouting dynamic control strategy; Using the three-zone digital twin model of the goaf, the real-time multi-source monitoring data analysis results and the real-time grouting dynamic control strategy are visualized, the real-time grouting dynamic control strategy is executed, and the data collection step continues.

2. The method for dynamic control of three-zone grouting in goaf based on reinforcement learning according to claim 1, characterized in that: A multi-source monitoring data analysis model was constructed using a deep learning algorithm. A grouting dynamic control model was constructed using a reinforcement learning algorithm. Furthermore, a three-zone digital twin model of the goaf was constructed using digital twin technology. The process involved the following steps: Use deep learning algorithms to build an initial multi-source monitoring data analysis model, and use reinforcement learning algorithms to build an initial grouting dynamic control model; Construct a three-dimensional simulation model of the three zones in the goaf. Combined with the initial multi-source monitoring data analysis model and the initial grouting dynamic control model, digital twin technology is used to construct a digital twin model of the three zones in the goaf. Collecting a number of historical multi-source monitoring data of the three zones of the goaf in the physical world corresponding to the digital twin model of the three zones of the goaf in the digital world, and preprocessing them to obtain a number of preprocessed historical multi-source monitoring data; Based on the three-dimensional simulation model of the three zones in the goaf and several pre-processed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic control model of the digital twin model of the three zones in the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic control model.

3. The method for dynamic control of three-zone grouting in goaf based on reinforcement learning according to claim 2, characterized in that: Using deep learning algorithms, we build an initial multi-source monitoring data analysis model. Using reinforcement learning algorithms, we build an initial grouting dynamic control model. This includes the following steps: Use the 3D-DBN-CNN-LSTM-ST-CNN algorithm to build an initial multi-source monitoring data analysis model; Using the MPO-MOGRPO-ICPO algorithm, an initial grouting dynamic control model is constructed; the initial grouting dynamic control model includes an initial meta-strategy optimization module, an initial control strategy generation module, and an initial control strategy optimization module; Set the objective function set, experience replay pool, actor network, and agent for the initial control strategy generation module, and use the initial network parameters of the actor network as the output parameters of the initial meta-strategy optimization module; The grouting dynamic control strategy generation problem is used as the simulation environment of the initial control strategy generation module, and the action space and state space are set for the intelligent agent of the initial control strategy generation module; Taking minimizing the strategy generation error as the optimization goal, the fitness function of the initial control strategy optimization module is defined, and the output of the Actor network of the initial control strategy generation module is used as the input of the initial control strategy optimization module.

4. The method for dynamic control of three-zone grouting in goaf based on reinforcement learning according to claim 3, characterized in that: Based on the three-dimensional simulation model of the goaf and several pre-processed historical multi-source monitoring data, the initial multi-source monitoring data analysis model and the initial grouting dynamic control model of the three-zone digital twin model of the goaf are trained to obtain the final multi-source monitoring data analysis model and the final grouting dynamic control model. The training process includes the following steps: The three-dimensional simulation model of the goaf is combined with some pre-processed historical multi-source monitoring data, and divided into model training set and model test set in a ratio of 7:

3. Based on the model training set, the initial multi-source monitoring data analysis model of the three-zone digital twin model of the goaf is trained to obtain an optimized multi-source monitoring data analysis model and generate several historical multi-source monitoring data analysis results; The optimized multi-source monitoring data analysis model is tested based on the model test set. If the test accuracy is greater than the accuracy threshold, the final multi-source monitoring data analysis model is output; otherwise, training continues. Traverse all the objective functions in the objective function set, train the initial grouting dynamic control model based on the analysis results of several historical multi-source monitoring data, obtain the final grouting dynamic control model, and store the historical grouting dynamic control experience generated during the training process in the experience replay pool.

5. The method for dynamic control of three-zone grouting in goaf based on reinforcement learning according to claim 4 is characterized in that: The historical multi-source monitoring data includes historical stratigraphic monitoring data of three zones in the goaf, historical environmental monitoring data, and historical equipment operation monitoring data; The real-time multi-source monitoring data includes real-time stratum monitoring data of three zones in the goaf, real-time environmental monitoring data and real-time equipment operation monitoring data.

6. The method for dynamic control of three-zone grouting in goaf based on reinforcement learning according to claim 5, characterized in that: Based on the digital twin model of the three zones in the goaf, real-time multi-source monitoring data of the three zones in the goaf is collected and spatially correlated to obtain the real-time multi-source monitoring data after spatial and temporal correlation. The process includes the following steps: Collect real-time multi-source monitoring data of the three zones of the goaf in the physical world, perform preprocessing, and input the obtained preprocessed real-time multi-source monitoring data into the digital twin model of the three zones of the goaf; Align the pre-processed real-time multi-source monitoring data on the timestamp to obtain the time-aligned real-time multi-source monitoring data; Perform spatial registration on the time-aligned real-time multi-source monitoring data to obtain spatially registered real-time multi-source monitoring data; Extract the real-time time features and real-time spatial features of the real-time multi-source monitoring data after spatial registration, and perform feature fusion on the real-time time features and real-time spatial features to obtain real-time fused spatiotemporal features; The real-time fusion of spatiotemporal features is spatiotemporally correlated in the three-zone digital twin model of the goaf to obtain real-time multi-source monitoring data after spatiotemporal correlation.

7. A reinforcement learning-based dynamic control system for three-zone grouting in goaf, used to implement the dynamic control method for three-zone grouting in goaf according to any one of claims 1 to 6, characterized in that: It includes a model building unit, a spatiotemporal correlation unit, a data analysis unit, a control strategy generation unit and a visualization unit that are connected in sequence.

Citation Information

Patent Citations

  • Coal mine comprehensive automatic management and control method and system based on digital twinborn technology

    CN118462314A

  • Multi-source information advanced grouting reinforcement intelligent device and effect evaluation method

    CN119777928A

  • Substation equipment fault early warning method and system

    CN120088959A