A method for constructing a large model for rock burst prediction based on multimodal data

Through the multimodal data fusion and Transformer architecture rock burst prediction model, the problem of insufficient adaptability and accuracy of the rock burst prediction model in different coal mining areas is solved, and the accurate prediction and prevention of rock burst risks are achieved.

CN119670963BActive Publication Date: 2025-09-19CHINA UNIV OF MINING & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411738841.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-09-19
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The existing rock burst prediction model has inaccurate positioning when identifying rock burst risk sources, low warning efficiency, and low generalization. It is difficult to adapt to the hazard level standards of different coal mines, and the model application in different mining areas is difficult.

Method used

By adopting multimodal data fusion technology, a large rock burst prediction model is constructed through sensor systems, mining information and geological structure data. The Transformer architecture is used for data processing, combined with the information entropy dynamic weight calculation method to achieve accurate prediction of rock burst risks.

Benefits of technology

The applicability and prediction accuracy of the rock burst prediction model in different mining areas have been improved, and accurate prediction and prevention of rock burst risks have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670963B_ABST
    Figure CN119670963B_ABST
Patent Text Reader

Abstract

A method for constructing a large-scale rock burst prediction model based on multimodal data, comprising the following steps: collecting data of different modalities to construct a multimodal dataset; preprocessing the multimodal dataset to construct a precursor pattern sequence for model training; converting the precursor pattern sequence into a corresponding grade form according to the characteristics of different mining areas, and assigning corresponding rock burst hazard level labels; using Transformer as the core framework to process the graded precursor pattern sequence, and ultimately achieving a prediction of the probability of rock burst level occurrence; utilizing a comprehensive index method to independently evaluate the degree of danger of mining information data and geological structure data, and combining the danger probability results output by the rock burst prediction module to comprehensively evaluate the overall danger level of the rock burst; the present invention can improve the applicability and prediction accuracy of the model under different mining conditions, and achieve accurate prediction of rock burst risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a rock burst model, in particular to a method for constructing a large rock burst prediction model based on multimodal data, and belongs to the technical field of coal mine monitoring and early warning. Background Art

[0002] Rock burst is a typical high-energy dynamic disaster in coal mining. It is characterized by suddenness and destructiveness, and can easily lead to serious consequences such as damage to mine equipment and personal injury. The mechanism of rock burst is complex and is affected by multiple factors such as the mine's geological structure, rock mass stress state, and mining depth. In recent years, with the continuous increase in the depth of mine resources, shallow resources have gradually become depleted, and the focus of underground mining activities has gradually shifted to deeper layers. Complex geological conditions have exacerbated the frequency of rock burst and the intensity of the disaster has also shown an upward trend. Therefore, how to achieve accurate monitoring and early warning of rock burst has become one of the core research issues in the field of mine safety.

[0003] Rock burst prediction involves multidisciplinary knowledge such as geology, rock mechanics, and data science. Its prediction accuracy and response timeliness are crucial to mine safety prevention and control. However, the existing monitoring and early warning systems still have shortcomings in the identification and prediction of rock burst risk sources. There are problems such as "inaccurate location of disaster sources and low early warning efficiency", making it difficult to accurately predict rock burst risks. In addition, the generalization of existing rock burst prediction models is low, and the hazard level standards of different coal mines are different, which makes it difficult to directly apply the constructed models to different mining areas. Traditional prediction methods mostly rely on single modal data or expert experience, or are often limited to a specific physical indicator. They lack adaptability when dealing with complex geological conditions, which seriously restricts the actual prevention and control effect of rock burst prediction models. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing a large-scale model for rock burst prediction based on multimodal data. By fusing the multimodal data of rock burst, the level and probability of large-energy events that may occur in the future are predicted in the time dimension. Combined with the information entropy dynamic weight calculation method based on time window design, the weight of the multimodal data is comprehensively evaluated to construct a basic large-scale model for rock burst prediction, improve the applicability and prediction accuracy of the model under different mining conditions, and realize accurate prediction of rock burst risks.

[0005] To achieve the above objectives, the present invention provides a method for constructing a large model for rock burst prediction based on multimodal data, comprising a multimodal data acquisition and preprocessing module, a rock burst prediction module, and a hazard level determination module. The specific steps are as follows:

[0006] S1. In the multimodal data acquisition and preprocessing module, first, a multimodal dataset is constructed by collecting data from different modalities. Then, the multimodal dataset is preprocessed to construct a precursor pattern sequence for model training. Finally, according to the characteristics of different mining areas, the precursor pattern sequence is converted into a corresponding grade form and the corresponding rock burst hazard level label is assigned.

[0007] S2. In the rock burst prediction module, the Transformer is used as the core framework to process the precursor pattern sequence after classification. The rock burst prediction module mainly includes the input embedding and position encoding layer, the Transformer encoder and the fully connected layer. These modules work together to ultimately predict the probability of rock burst level occurrence.

[0008] S3. In the hazard level determination module, the comprehensive index method is used to independently evaluate the hazard levels of mining information data and geological structure data, and the overall hazard level of rock burst is comprehensively evaluated in combination with the hazard probability results output by the rock burst prediction module. Specifically, for the three parts of data, namely mining information, geological structure, and prediction results, their contribution ratios in the comprehensive index are allocated through a weighting method, thereby achieving a comprehensive prediction of the hazard level.

[0009] The multimodal data set in step S1 of the present invention includes dynamic data consisting of sensor system data and mining information data and static data consisting of geological structure data;

[0010] The sensor system data is collected in real time through high-precision sensors rationally arranged in the mine, mainly including microseismic monitoring waveform data, ground sound waveform data, rock stress waveform data and electromagnetic signal waveform data. The microseismic monitoring waveform data is collected through the microseismic sensor array arranged in the mine to capture the vibration signal caused by the stress change of the rock mass. The ground sound waveform data is collected through the ground sound sensors arranged in the mine to capture the tiny sound fluctuations of the rock mass, reflecting the dynamic changes of the stress of the rock stratum. The rock stress waveform data is collected through the stress sensor to collect the dynamic changes of the stress in the rock stratum in the mine. The electromagnetic signal waveform data is collected through the electromagnetic sensors placed in the mine to monitor the changes of electromagnetic signals in the rock stratum during the stress process in real time. The joint application of multiple sensing systems realizes the high-frequency collection of multi-dimensional data, providing multi-angle information for rock burst prediction.

[0011] Mining information data is used to describe the mining status of the mine. As the mining process continues to change, it is crucial for the risk assessment and prediction of rock burst. Mining information data includes the closest distance W between the mining location and irregular working surfaces such as "knife handle" shapes or multiple working surfaces, and areas where the cuts and stop lines are not aligned. e 1, the shortest distance W between the mining position and the "square" area of ​​the goaf of the working face e 2 , the shortest distance W between the mining location and the intersection area of ​​the "triangle" roadway e 3 , mining speed W e 4 And the closest distance between the mining location and the structural features around the mine, such as the closest distance W between the mining location and the fault (drop greater than 3m) e 5 , the closest distance to the fold (inclination greater than 15°) W e 6 and the shortest distance to the goaf W e 7 , the change rate of coal seam thickness (relative to the average coal thickness) at the mining location W e 8 ;

[0012] Geological structure data is used to describe the geological factors of mining mines and evaluate the overall rock burst hazard level of coal mines before mining. Geological structure data includes geological data and mining data. Among them, geological data includes the historical number of rock burst occurrences W1 1 , mining depth W1 2 , the distance between the hard and thick rock layer in the overlying fracture zone and the coal seam W1 3 , roof rock thickness characteristic parameter W1 4 , the degree of structural stress concentration in the mining area W1 5 , uniaxial compressive strength of coal W1 6 and the elastic energy index W1 of coal 7 ; Mining data includes the degree of pressure relief of the protective layer W2 1 , the horizontal distance W2 between the working face and the coal pillar left by mining the upper protective layer 2 , Relationship with adjacent goaf W2 3 , working surface strength W2 4 , Section coal pillar width W2 5 , bottom coal thickness W2 6 , the distance from the goaf when excavating into the goaf is W2 7 , the distance from the goaf when advancing to the goaf W2 8 , distance from the fault W2 9 , distance from the fold W2 10 and the distance W2 from the coal seam phase change zone 11 ;

[0013] The specific method of step S1 of the present invention is as follows:

[0014] S1.1: First, perform data preprocessing on the raw data of the sensor system. For the microseismic monitoring waveform data and ground sound waveform data, use the bandpass filtering method to remove low-frequency or high-frequency background noise;

[0015] For rock stress waveform data, outlier detection methods are used to remove data deviations caused by sensor errors or environmental interference;

[0016] For electromagnetic signal data, wavelet transform is used to perform denoising and extract effective electromagnetic signal components;

[0017] S1.2: Secondly, the denoised sensor raw data is converted into a format. Specifically, the microseismic monitoring waveform data and ground sound waveform data are converted into time-energy format data.

[0018] Convert rock mass stress waveform data into time-stress format data;

[0019] Convert electromagnetic signal waveform data into time-magnetic field data;

[0020] Through this step, high-quality multimodal data that is suitable for model training and prediction needs can be generated, providing reliable input support for subsequent modeling and analysis. Therefore, the sensor system dataset d i Can be recorded as Then the j-th data of the i-th sensor can be expressed as:

[0021]

[0022] Where: d i represents the i-th sensor system dataset;

[0023] T i j It is represented as the time corresponding to the jth data of the i-th sensor;

[0024] It is represented as the energy / stress / magnetic field corresponding to the jth data of the i-th sensor;

[0025] Use k fixed time windows to count sensor system data, the number of data is n, then the i-th sensor system time window sequence data set is It can be expressed as:

[0026]

[0027] Statistical analysis of time window series datasets The obtained sensor data set is recorded as U, and the k-th time window data of the i-th sensor is recorded as It can be expressed as:

[0028]

[0029] Where: is the number of the kth time window of the i-th sensor;

[0030] is the maximum energy / stress / magnetic field in the kth time window;

[0031] is the average energy / stress / magnetic field in the kth time window;

[0032] f i k is the frequency of energy / stress / magnetic field in the kth time window;

[0033] According to the sensor data set U, the precursor pattern sequence w is constructed, and the precursor pattern sequence of the i-th sensor e can be expressed as

[0034]

[0035] Where: g is the sampling step length;

[0036] p is the length of the precursor pattern sequence, and the precursor pattern sequence set W of the constructed sensor i i Expressed as:

[0037]

[0038] Where: q is the total number of precursor pattern sequences;

[0039] S1.3: The sensor data in the precursory pattern sequence are standardized. By converting numerical data such as microseismic energy / magnetic field / stress, and frequency into graded information, the grade is used instead of the specific numerical value as the model input, thereby effectively improving the adaptability and prediction performance of the model under different mining conditions.

[0040] The input embedding and position encoding layer of step S2 of the present invention is specifically:

[0041] Input embedding maps the input sequence (precursor pattern sequence) to a high-dimensional space, forming a fixed length suitable for model processing. For each input segment x i , the embedded vector e is obtained by linear transformation i :

[0042] e i =W e x i +b e (6)

[0043] Where: W eand b e are the weights and biases of the embedding layer;

[0044] Since the Transformer itself does not have the ability to process position information, time series information is introduced through the position encoding layer. The position encoding layer uses sine and cosine functions to generate position information:

[0045]

[0046]

[0047] Where: pos represents the sequence position;

[0048] a is the dimension index;

[0049] d model is the embedding dimension;

[0050] The sequence obtained by adding the positional encoding layer to the input embedding can be represented as:

[0051] z0=[e1+PE1,e2+PE2,...,e N +PE N ] (9)

[0052] The Transformer encoder is specifically:

[0053] The Transformer encoder is the core of the entire network, used to extract global features of the precursor pattern sequence. The module is composed of multiple encoders stacked together. Each encoder includes a multi-head self-attention mechanism, a feedforward neural network, and a residual connection and normalization. The output of the encoder is a high-dimensional feature representation that contains complex relationships between different time segments. Specifically:

[0054] The multi-head self-attention mechanism is used to calculate the display weights between sequence segments, thereby dynamically capturing the temporal dependency and cross-modal correlation of rock burst precursor patterns and effectively mining potential feature patterns. The multi-head self-attention mechanism includes a self-attention mechanism and a multi-head mechanism.

[0055] The self-attention mechanism generates a query vector, a key vector, and a value vector for each input sequence z0, and calculates the similarity weight through the dot product operation (MatMul):

[0056] Q=z0W Q ,K=z0W K ,V=z0W V (10)

[0057] Where: W Q , W K , W V is a learnable weight matrix;

[0058] The similarity weights are scaled by the dot product operation and then normalized by the softmax activation function:

[0059]

[0060] Where: d k is the dimension of the key vector;

[0061] In order to enhance the feature extraction capability of the model, multiple heads are used to calculate attention in parallel, and each head has an independent W Q , W K , W V :

[0062] MultiHead(Q,K,V)=Concat(head1,...,head h )W O (12)

[0063] Where:

[0064]

[0065] h is the number of attention heads, W O is the output linear transformation matrix;

[0066] The feedforward neural network performs a nonlinear transformation on the features output by the self-attention mechanism. After the multi-head attention mechanism, the feature vector at each position passes through two layers of fully connected networks and adds a ReLU activation function:

[0067] FFN(x)=W2(ReLU(W1x+b1))+b2 (14)

[0068] Where: W1 and b1 are the weight and bias of the first fully connected network layer;

[0069] W2 and b2 are the weights and biases of the second fully connected network layer;

[0070] Residual connections and normalization

[0071] To avoid the problems of gradient disappearance and gradient explosion, residual connections and layer normalization are added after each sublayer:

[0072] Output=LayerNorm(x+SubLayer(x)) (15)

[0073] Where: LayerNorm(·) is the layer normalization calculation;

[0074] SubLayer(x) is the output of a multi-head attention or feedforward neural network;

[0075] The fully connected layer is specifically:

[0076] The high-dimensional feature representation Z generated by the Transformer encoder L As the input of the fully connected layer, after one or more layers of linear transformation, the final output is the probability distribution of the rock burst danger level, and the Softmax activation function can be used to output the level with the highest probability as the prediction result:

[0077] p c =softmax(W d (x)+b d ) (16)

[0078] Where: W d and b d are the weights and biases of the first fully connected layer;

[0079] p c is the predicted probability that the rock burst category is predicted to be c.

[0080] The details of step S3 of the present invention are as follows:

[0081] First, the rock burst hazard level (RL) is normalized to the [0, 1] interval. The mining information, geological structure, and prediction results data are also divided into the [0, 1] interval to facilitate the final assessment of the rock burst hazard.

[0082] The mining information data is divided into criteria using the comprehensive index method, in which the specific criteria for classification of some factors can be modified accordingly according to actual conditions;

[0083] Among them, the mining information data impact factor W e The calculation is as follows:

[0084]

[0085] The geological structure data is divided into rules using the comprehensive index method, and the geological structure analysis of the impact of the comprehensive index method on the above geological data and mining data is carried out:

[0086] Geological data: Mining data:

[0087] Geological structure data impact factor W g Select the maximum value of geological data and mining comprehensive index, namely:

[0088] W g =max{W g1 , W g2} (19)

[0089] Deep learning data impact factor W m To output the danger level value, we first determine the danger level (none, weak, medium, strong) by the maximum probability value output by the model, and then divide the maximum probability value into 5 levels. At the same time, we divide the 4 danger level ranges into 5 levels corresponding to the maximum probability value range, which is used as the deep learning data influencing factor W. m , further improving the accuracy and practicality of the prediction.

[0090] Analysis shows that the maximum value of the probability of the model outputting the danger level (p c ) max The range is (0.25, 1]. In order to fully reflect the different levels of danger, the present invention constructs different deep learning data impact factors W based on the distribution characteristics of the model output probability. m ,Through this classification method, each hazard level can not only reflect the ,probabilistic output results of the model, but also effectively improve the ,accuracy of hazard classification, thereby achieving a more reliable risk assessment of ,rock burst.

[0091] Compared with the existing technology, the present invention uses a multimodal data acquisition and preprocessing module and multimodal data fusion technology to convert the raw data collected by the sensor system into a precursor pattern sequence. Compared with the existing method of directly using raw data, this method innovatively uses a hierarchical form to standardize the data, which can significantly improve the adaptability and prediction accuracy of the model under different mining conditions. In the rock burst prediction module, a Transformer-based model architecture is used. Unlike the traditional deep learning method that directly outputs fixed results, the present invention uses a probability distribution form to fine-tune the probability of rock burst hazard level. In the hazard level determination module, a dynamic weight calculation method based on time window information entropy is proposed to achieve multi-source information fusion of mining information, geological structure data, and prediction results to comprehensively assess the hazard level of rock burst. The present invention provides a method for constructing a large rock burst prediction model based on multimodal data. After training the basic large model on historical data of other working faces, it can be transferred and applied to new working faces, providing a reference for the time series prediction and prevention of rock burst, improving the applicability and prediction accuracy of the model under different mining conditions, and achieving accurate prediction of rock burst risk. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 This is a schematic diagram of the overall architecture of the present invention;

[0093] Figure 2 Schematic diagram of the self-attention mechanism of the present invention;

[0094] Figure 3 Schematic diagram of the multi-head attention mechanism of the present invention;

[0095] Figure 4 Schematic diagram of residual connection and normalization of the present invention;

[0096] Figure 5 Schematic diagram of a feedforward neural network of the present invention;

[0097] Figure 6 Schematic diagram of residual connection and layer normalization of the present invention. DETAILED DESCRIPTION

[0098] The present invention will be further described below with reference to the accompanying drawings.

[0099] like Figure 1 As shown, a method for constructing a large-scale rock burst prediction model based on multimodal data includes a multimodal data acquisition and preprocessing module, a rock burst prediction module, and a hazard level determination module. The specific steps are as follows:

[0100] S1. In the multimodal data acquisition and preprocessing module, first, a multimodal dataset is constructed by collecting data from different modalities. Then, the multimodal dataset is preprocessed to construct a precursor pattern sequence for model training. Finally, according to the characteristics of different mining areas, the precursor pattern sequence is converted into a corresponding grade form and the corresponding rock burst hazard level label is assigned.

[0101] The multimodal data set is composed of dynamic data consisting of sensor system data and mining information data, and static data consisting of geological structure data.

[0102] The sensor system data is collected in real time through high-precision sensors rationally arranged in the mine, and mainly includes microseismic monitoring waveform data, ground sound waveform data, rock stress waveform data, and electromagnetic signal waveform data; the microseismic monitoring waveform data captures the vibration signal caused by the stress change of the rock mass through the microseismic sensor array arranged in the mine; the ground sound waveform data captures the tiny sound fluctuations of the rock mass through the ground sound sensors arranged in the mine, reflecting the dynamic changes of the stress of the rock stratum; the rock stress waveform data collects the dynamic changes of the stress in the rock stratum in the mine through the stress sensor; the electromagnetic signal waveform data monitors the changes of the electromagnetic signal in the process of the rock stratum being stressed in real time through the electromagnetic sensors placed in the mine; through the joint application of multiple sensing systems, high-frequency collection of multi-dimensional data is achieved, providing multi-angle information for rock burst prediction;

[0103] Mining information data is used to describe the mining status of the mine. As the mining process continues to change, it is crucial for the risk assessment and prediction of rock burst. Mining information data includes the closest distance W between the mining location and irregular working surfaces such as "knife handle" shapes or multiple working surfaces, and areas where the cuts and stop lines are not aligned. e1 , the shortest distance W between the mining position and the "square" area of ​​the goaf of the working face e 2 , the shortest distance W between the mining location and the intersection area of ​​the "triangle" roadway e 3 , mining speed W e 4 And the closest distance between the mining location and the structural features around the mine, such as the closest distance W between the mining location and the fault (drop greater than 3m) e 5 , the closest distance to the fold (inclination greater than 15°) W e 6 and the shortest distance to the goaf W e 7 , the change rate of coal seam thickness (relative to the average coal thickness) at the mining location W e 8 Mining information data can provide the basis for the changes in the overall stress field of the mine and is the key part of the multi-modal data input for building the rock burst prediction model;

[0104] Geological structure data is used to describe the geological factors of mining mines and evaluate the overall rock burst hazard level of coal mines before mining. Geological structure data includes geological data and mining data. Among them, geological data includes the historical number of rock burst occurrences W1 1 , mining depth W1 2 , the distance between the hard and thick rock layer in the overlying fracture zone and the coal seam W1 3 , roof rock thickness characteristic parameter W1 4 , the degree of structural stress concentration in the mining area W1 5 , uniaxial compressive strength of coal W1 6 and the elastic energy index W1 of coal 7 ; Mining data includes the degree of pressure relief of the protective layer W2 1 , the horizontal distance W2 between the working face and the coal pillar left by mining the upper protective layer 2 , Relationship with adjacent goaf W2 3 , working surface strength W2 4 , Section coal pillar width W2 5 , bottom coal thickness W2 6 , the distance from the goaf when excavating into the goaf is W2 7 , the distance from the goaf when advancing to the goaf W2 8 , distance from the fault W2 9 , distance from the fold W2 10 and the distance W2 from the coal seam phase change zone 11 ;

[0105] The specific method of step S1 is as follows:

[0106] S1.1: First, due to the high noise interference in the mine environment, to ensure high-quality multimodal data when input into the model, the raw data of the sensor system data needs to be preprocessed. For the microseismic monitoring waveform data and ground sound waveform data, the low-frequency and high-frequency background noise are removed using the bandpass filtering method.

[0107] For rock stress waveform data, outlier detection methods are used to remove data deviations caused by sensor errors or environmental interference;

[0108] For electromagnetic signal data, wavelet transform is used to perform denoising and extract effective electromagnetic signal components;

[0109] S1.2: Secondly, to ensure that the denoised sensor system data is consistent and suitable for the training and prediction process of the rock burst prediction model, the denoised sensor raw data needs to be formatted. Specifically, the microseismic monitoring waveform data and ground sound waveform data need to be converted into time-energy format data.

[0110] Convert rock mass stress waveform data into time-stress format data;

[0111] Convert electromagnetic signal waveform data into time-magnetic field data;

[0112] Through this step, high-quality multimodal data that is suitable for model training and prediction needs can be generated, providing reliable input support for subsequent modeling and analysis. Therefore, the sensor system dataset d i Can be recorded as Then the j-th data of the i-th sensor can be expressed as:

[0113]

[0114] Where: d i represents the i-th sensor system dataset;

[0115] T i j It is represented as the time corresponding to the jth data of the i-th sensor;

[0116] It is represented as the energy / stress / magnetic field corresponding to the jth data of the i-th sensor;

[0117] Use k fixed time windows to count sensor system data, the number of data is n, then the i-th sensor system time window sequence data set is It can be expressed as:

[0118]

[0119] Statistical analysis of time window series datasets The obtained sensor data set is recorded as U, and the k-th time window data of the i-th sensor is recorded as It can be expressed as:

[0120]

[0121] Where: is the number of the kth time window of the i-th sensor;

[0122] is the maximum energy / stress / magnetic field in the kth time window;

[0123] is the average energy / stress / magnetic field in the kth time window;

[0124] f i k is the frequency of energy / stress / magnetic field in the kth time window;

[0125] According to the sensor data set U, the precursor pattern sequence w is constructed, and the precursor pattern sequence of the i-th sensor e can be expressed as

[0126]

[0127] Where: g is the sampling step length;

[0128] p is the length of the precursor pattern sequence, and the precursor pattern sequence set W of the constructed sensor i i like Figure 2 As shown, it can be expressed as:

[0129]

[0130] Where: q is the total number of precursor pattern sequences;

[0131] S1.3: Given the varying degrees of danger posed by different mining areas under the same microseismic energy, magnetic field, stress, or frequency, directly inputting raw data into the model can easily lead to the model being unable to adapt to the specific conditions of each mining area, exhibiting insufficient generalization. To address this issue, this method standardizes the sensor data in the precursory pattern sequence. By converting numerical data such as microseismic energy, magnetic field, stress, and frequency into graded information, the model uses these grades instead of specific values ​​as input, effectively improving the model's adaptability and predictive performance under different mining conditions.

[0132] Taking maximum microseismic energy and frequency as an example, different classifications can be applied to different coal mine conditions based on specific needs. Table 1 shows examples of microseismic energy and frequency classifications in two coal mines. For unmined coal mines, initial classification standards can be developed through statistical analysis of historical data from other working faces in the mine. Appropriate adjustments can be made to the classification standards once sufficient data is accumulated.

[0133] Table 1 Information on classification of different coal mines

[0134]

[0135] In addition, the definition of hazard levels may vary among coal mines. For example, as shown in Table 2, different hazard level labels need to be set according to the actual situation of the coal mine and used as classification labels in subsequent model training to improve the prediction accuracy of the model in diverse application scenarios.

[0136] Table 2 Classification of different coal mine hazard level labels

[0137]

[0138] S2. In the rock burst prediction module, the Transformer is used as the core framework to process the precursor pattern sequence after classification. The rock burst prediction module mainly includes the input embedding and position encoding layer, the Transformer encoder and the fully connected layer. These modules work together to ultimately predict the probability of rock burst level occurrence.

[0139] S2.1: Input embedding and position encoding layers are specifically:

[0140] Input embedding maps the input sequence (precursor pattern sequence) to a high-dimensional space, forming a fixed length suitable for model processing. For each input segment x i , the embedded vector e is obtained by linear transformation i :

[0141] e i =W e x i +b e (6)

[0142] Where: W e and b e are the weights and biases of the embedding layer;

[0143] Since the Transformer itself does not have the ability to process position information, time series information is introduced through the position encoding layer. The position encoding layer uses sine and cosine functions to generate position information:

[0144]

[0145] Where: pos represents the sequence position;

[0146] a is the dimension index;

[0147] d model is the embedding dimension;

[0148] The sequence obtained by adding the positional encoding layer to the input embedding can be represented as:

[0149] z0=[e1+PE1,e2+PE2,...,e N +PE N ] (9)

[0150] S2.2: Transformer encoder is specifically:

[0151] The Transformer encoder is the core of the entire network, used to extract global features of the precursor pattern sequence. The module is composed of multiple stacked encoders. Each encoder includes a multi-head self-attention mechanism, a feedforward neural network, a residual connection, and normalization. The encoder output is a high-dimensional feature representation that contains complex relationships between different time segments.

[0152] S2.2.1: Multi-head self-attention mechanism

[0153] Multi-head self-attention mechanism such as Figure 3 、 Figure 4 As shown, it is used to calculate the display weights between sequence segments, thereby dynamically capturing the temporal dependence and cross-modal correlation of rock burst precursor patterns and effectively mining potential feature patterns. The multi-head attention mechanism includes self-attention mechanism and multi-head mechanism;

[0154] The self-attention mechanism generates a query vector, a key vector, and a value vector for each input sequence z0, and calculates the similarity weight through the dot product operation (MatMul):

[0155] Q=z0W Q ,K=z0W K ,V=z0W V (10)

[0156] Where: W Q , W K , W V is a learnable weight matrix;

[0157] The similarity weights are scaled by the dot product operation and then normalized by softmax:

[0158]

[0159] Where: dk is the dimension of the key vector, which is used to prevent gradient instability caused by excessively large dot product values.

[0160] In order to enhance the feature extraction capability of the model, multiple heads are used to calculate attention in parallel, and each head has an independent W Q , W K , W V :

[0161] MultiHead(Q,K,V)=Concat(head1,...,head h )W O (12)

[0162] Where:

[0163]

[0164] h is the number of attention heads, W O is the output linear transformation matrix;

[0165] S2.2.2: Feedforward Neural Networks

[0166] Feedforward neural networks such as Figure 5 As shown in the figure, the features output by the self-attention mechanism are transformed nonlinearly to further improve the expressiveness of the model. After the multi-head attention mechanism, the feature vector of each position is passed through two layers of fully connected networks and the ReLU activation function is added:

[0167] FFN(x)=W2(ReLU(W1x+b1))+b2 (14)

[0168] Where: W1 and b1 are the weight and bias of the first fully connected network layer;

[0169] W2 and b2 are the weights and biases of the second fully connected network layer;

[0170] S2.2.3: Residual Connections and Normalization

[0171] To avoid the problem of gradient disappearance and gradient explosion, residual connection and layer normalization are added after each sublayer, such as Figure 6 As shown:

[0172] Output=LayerNorm(x+SubLayer(x)) (15)

[0173] Where: LayerNorm(·) is the layer normalization calculation;

[0174] SubLayer(x) is the output of a multi-head attention or feedforward neural network;

[0175] S2.3: The fully connected layer is specifically:

[0176] The high-dimensional feature representation Z generated by the Transformer encoder L As the input of the fully connected layer, after one or more layers of linear transformation, the final output is the probability distribution of the rock burst danger level, and the Softmax activation function can be used to output the level with the highest probability as the prediction result:

[0177] p c =softmax(W d (x)+b d ) (16)

[0178] Where: W d and b d are the weights and biases of the first fully connected layer;

[0179] p c is the predicted probability that the rock burst category is predicted to be c.

[0180] S3. In the hazard level determination module, a comprehensive index method is used to independently evaluate the hazard level of mining information and geological structure data. Combined with the hazard probability results output by the rock burst prediction module, an information entropy weighting method based on a time window design is used to comprehensively assess the overall hazard level of rock burst. Specifically, a weighting method is used to assign the contribution of the three data components (mining information, geological structure, and prediction results) to the hazard level, thereby achieving a comprehensive prediction of the hazard level.

[0181] First, the rock burst hazard level RL is normalized to the interval [0, 1], as shown in Table 3. The mining information, geological structure, and prediction results are also divided into the interval [0, 1] to facilitate the final assessment of the rock burst hazard;

[0182] Table 3 Rock burst hazard levels

[0183] Hazard Level Danger level RL corresponding range none 0≤RL<0.25 weak 0.25≤RL<0.5 middle 0.5≤RL<0.75 powerful 0.75≤RL<1

[0184] S3.1: The classification criteria for mining information data using the comprehensive index method are shown in Table 4. The specific criteria for classification of some factors can be modified accordingly based on actual conditions;

[0185] Table 4 Mining information classification criteria

[0186]

[0187]

[0188] Mining information data impact factor W e The calculation is as follows:

[0189]

[0190] S3.2: The classification rules of geological structure data using the comprehensive index method are shown in Table 5;

[0191] Table 5 Geological structure division criteria (a) Geological structure division criteria under the influence of geological data

[0192]

[0193]

[0194] (b) Geological structure classification criteria under the influence of mining data

[0195]

[0196]

[0197] The geological structure analysis of the influence of the above geological data and mining data using the comprehensive index method:

[0198] Geological data: Mining data:

[0199] Geological structure data impact factor W g Select the maximum value of geological data and mining comprehensive index, namely:

[0200] W g =max{W g1 , W g2} (19)

[0201] S3.3: Traditional deep learning models usually use the maximum value of the model output category probability as the prediction result. When the maximum probability value is high, the model's credibility in its output result is relatively high; however, when the probability values ​​of multiple categories are close, the model's judgment on category attribution may be uncertain, thereby reducing the reliability of the prediction result. To address this problem, the present invention proposes a method that comprehensively considers the maximum probability value of the model output and the impact ground pressure hazard level. First, the corresponding hazard level (none, weak, medium, strong) is determined based on the maximum probability value output by the model, and the maximum probability value is further divided into five sub-level ranges; then, the range of the determined hazard level (as shown in Table 3) is divided into 5 refined hazard degree values ​​corresponding to the five sub-level ranges of the maximum probability value; finally, based on the range of the maximum probability value, the hazard degree value is determined as the deep learning data influencing factor W m , in order to improve the accuracy and practicality of the prediction.

[0202] Therefore, the present invention proposes a method that comprehensively considers the maximum probability of the model output and the rock burst hazard level. First, the hazard level (none, weak, medium, strong) is determined by the maximum probability value output by the model. Then, the maximum probability value is divided into 5 levels. At the same time, the 4 hazard level ranges are divided into 5 levels corresponding to the maximum probability value range, thereby outputting the hazard level value as the deep learning data influencing factor W. m , further improving the accuracy and practicality of the prediction.

[0203] Analysis shows that the maximum value of the probability of the model outputting the danger level (p c ) max The range is (0.25, 1]. In order to fully reflect the different levels of danger, the present invention constructs different deep learning data impact factors W based on the distribution characteristics of the model output probability. m Specific classification criteria are detailed in Table 6. Through this classification method, each hazard level can not only reflect the probability output results of the model, but also effectively improve the classification accuracy of the hazard level, thereby achieving a more reliable risk assessment of rock burst.

[0204] Table 6 Output criteria for deep learning data impact factors

[0205]

[0206] In order to further comprehensively evaluate the danger level of rock burst, the present invention proposes a method for calculating information entropy weight based on time window design. e , geological structure data impact factor W g , Deep Learning Data Impact Factor W m , use weight division to determine the weight α of each part e , α g , α m , satisfying α e +α g +α m =1.

[0207] First, considering that the weights should change dynamically with the mining process, we use the same time window as the precursor pattern sequence to count these three types of data and calculate the probability distribution of each type of data:

[0208]

[0209] Among them, W k (l) represents the mining information data impact factor, geological structure data impact factor, and deep learning data impact factor of the lth sample, and b represents the total number of samples.

[0210] Then, calculate the information entropy of each type of data:

[0211]

[0212] Among them E k represents the information entropy of the kth category of data, and ln(b) is the normalized coefficient of entropy.

[0213] Therefore, the weight calculation formula for each part is:

[0214]

[0215] Among them, α e , α g , α m The impact factor W of the information data e , geological structure data impact factor W g , Deep Learning Data Impact Factor W m The weight of .

[0216] Finally, the risk level RL for the predicted time period is calculated as:

[0217] RL=α e W e +α g W g +α m W m (twenty three)

[0218] Table 3 is used to determine the degree of danger (none, weak, moderate, strong) in the forecast period to predict the rock burst danger.

Claims

1. A method for constructing a large rock burst prediction model based on multimodal data, characterized in that: It includes a multimodal data acquisition and preprocessing module, a rock burst prediction module, and a hazard level determination module. The specific steps are as follows: S1. In the multimodal data acquisition and preprocessing module, first, a multimodal dataset is constructed by collecting data from different modalities. Then, the multimodal dataset is preprocessed to construct a precursor pattern sequence for model training. Finally, according to the characteristics of different mining areas, the precursor pattern sequence is converted into a corresponding grade form and the corresponding rock burst hazard level label is assigned. S2. In the rock burst prediction module, the Transformer is used as the core framework to process the precursor pattern sequence after classification. The rock burst prediction module mainly includes the input embedding and position encoding layer, the Transformer encoder and the fully connected layer. These modules work together to ultimately predict the probability of rock burst level occurrence. S3. In the hazard level determination module, the comprehensive index method is used to independently evaluate the hazard levels of mining information data and geological structure data, and the overall hazard level of rock burst is comprehensively evaluated in combination with the hazard probability results output by the rock burst prediction module. Specifically, for the three parts of data, namely mining information, geological structure, and prediction results, their contribution ratios in the comprehensive index are allocated through a weighting method, thereby achieving a comprehensive prediction of the hazard level.

2. The method for constructing a large-scale rock burst prediction model based on multimodal data according to claim 1, characterized in that: The multimodal data set in step S1 is formed by fusion of dynamic data consisting of sensor system data and mining information data and static data consisting of geological structure data; The sensor system data is collected in real time through high-precision sensors rationally arranged throughout the mine, including microseismic monitoring waveform data, geoacoustic waveform data, rock stress waveform data, and electromagnetic signal waveform data. The microseismic monitoring waveform data is collected through the microseismic sensor array arranged throughout the mine, capturing vibration signals caused by changes in rock stress. The geoacoustic waveform data is collected through geoacoustic sensors arranged throughout the mine, capturing minute sound fluctuations in the rock mass, reflecting the dynamic changes in rock stratum stress. Rock stress waveform data is collected through stress sensors to monitor the dynamic changes in stress in the rock formations of the mine. Electromagnetic signal waveform data is collected through electromagnetic sensors installed in the mine to monitor the changes in electromagnetic signals during the stress process of the rock formations in real time. Through the combined application of multiple sensing systems, high-frequency collection of multi-dimensional data is achieved, providing multi-angle information for rock burst prediction. Mining information data is used to describe the mining status of the mine. As the mining process continues to change, it is crucial for the risk assessment and prediction of rock burst. Mining information data includes the closest distance between the mining location and the "knife handle" irregular working face or multiple working faces, the opening of the cutting hole, and the area where the stop line is not aligned. The shortest distance between the mining position and the "square" area of ​​the goaf of the working face The shortest distance between the mining location and the intersection area of ​​the "triangle" roadway Mining speed and the closest distance between the mining location and the structural features surrounding the mine, including the closest distance between the mining location and the fault Closest distance to the fold and the closest distance to the goaf Coal seam thickness change rate at mining location Geological structure data is used to describe the geological factors of mining mines and evaluate the overall rock burst hazard level of coal mines before mining. Geological structure data includes geological data and mining data. Among them, geological data includes the historical number of rock burst occurrences W1 1 , mining depth W1 2 , the distance between the hard and thick rock layer in the overlying fracture zone and the coal seam W1 3 , roof rock thickness characteristic parameter W1 4 , the degree of structural stress concentration in the mining area W1 5 , uniaxial compressive strength of coal W1 6 and the elastic energy index W1 of coal 7 ; Mining data includes the degree of pressure relief of the protective layer W2 1 , the horizontal distance between the working face and the coal pillar left by mining the upper protective layer Relationship with adjacent goaf Working surface strength Sectional coal pillar width Thickness of bottom coal The distance from the goaf when excavating into it The distance from the goaf when advancing towards it Distance from the fault Distance from folds and the distance from the coal seam phase change zone 3. The method for constructing a large-scale rock burst prediction model based on multimodal data according to claim 2, characterized in that: The specific method of step S1 is as follows: S1.1: First, perform data preprocessing on the raw data of the sensor system. For the microseismic monitoring waveform data and ground sound waveform data, use the bandpass filtering method to remove low-frequency or high-frequency background noise; For rock stress waveform data, outlier detection methods are used to remove data deviations caused by sensor errors or environmental interference; For electromagnetic signal data, wavelet transform is used to perform denoising and extract effective electromagnetic signal components; S1.2: Secondly, the denoised sensor raw data is converted into a format. Specifically, the microseismic monitoring waveform data and ground sound waveform data are converted into time-energy format data. Convert rock mass stress waveform data into time-stress format data; Convert electromagnetic signal waveform data into time-magnetic field data; Through this step, high-quality multimodal data that is suitable for model training and prediction needs can be generated, providing reliable input support for subsequent modeling and analysis. Therefore, the sensor system dataset d i Recorded as Then the j-th data of the i-th sensor is expressed as: Where: d i represents the i-th sensor system dataset; T i j It is represented as the time corresponding to the jth data of the i-th sensor; It is represented as the energy / stress / magnetic field corresponding to the jth data of the i-th sensor; Use k fixed time windows to count sensor system data, the number of data is n, then the i-th sensor system time window sequence data set is Expressed as: Statistical analysis of time window series datasets The obtained sensor data set is recorded as U, and the k-th time window data of the i-th sensor is recorded as Expressed as: Where: is the number of the kth time window of the i-th sensor; is the maximum energy / stress / magnetic field in the kth time window; is the average energy / stress / magnetic field in the kth time window; f i k is the frequency of energy / stress / magnetic field in the kth time window; According to the sensor data set U, the precursor pattern sequence w is constructed, and the precursor pattern sequence of the i-th sensor e is expressed as Where: g is the sampling step length; p is the length of the precursor pattern sequence, and the precursor pattern sequence set W of the constructed sensor i i Expressed as: Where: q is the total number of precursor pattern sequences; S1.3: The sensor data in the precursory pattern sequence were standardized. By converting the numerical data of microseismic energy / magnetic field / stress and frequency into graded information, the grades were used instead of specific values ​​as model input, thereby effectively improving the adaptability and prediction performance of the model under different mining conditions.

4. The method for constructing a large-scale rock burst prediction model based on multimodal data according to claim 2, characterized in that: The input embedding and position encoding layers of step S2 are specifically: Input embedding maps the precursor pattern sequence to a high-dimensional space, forming a fixed length suitable for model processing. For each input segment x i , the embedded vector e is obtained by linear transformation i : e i =W e x i +b e (6) Where: W e and b e are the weights and biases of the embedding layer; Since the Transformer itself does not have the ability to process position information, time series information is introduced through the position encoding layer. The position encoding layer uses sine and cosine functions to generate position information: Where: pos represents the sequence position; a is the dimension index; d model is the embedding dimension; The sequence representation obtained after the input embedding plus the positional encoding layer is: z0=[e1+For1,e2+For2,…,and N +PE N ] (9) The Transformer encoder is specifically: The Transformer encoder is the core of the entire network, used to extract global features of the precursor pattern sequence. The module is composed of multiple encoders stacked together. Each encoder includes a multi-head self-attention mechanism, a feedforward neural network, and a residual connection and normalization. The output of the encoder is a high-dimensional feature representation that contains complex relationships between different time segments. Specifically: The multi-head self-attention mechanism is used to calculate the display weights between sequence segments, thereby dynamically capturing the temporal dependency and cross-modal correlation of rockburst precursor patterns. The multi-head self-attention mechanism includes a self-attention mechanism and a multi-head mechanism. The self-attention mechanism generates a query vector, a key vector, and a value vector for each input sequence z0, and calculates the similarity weight through the dot product operation: Q=z0W Q ,K=z0W K ,V=z0W V (10) Where: W Q , W K , W V is a learnable weight matrix; The similarity weights are scaled by the dot product operation and then normalized by the softmax activation function: Where: d k is the dimension of the key vector; In order to enhance the feature extraction capability of the model, multiple heads are used to calculate attention in parallel, and each head has an independent W Q , W K , W V : MultiHead(Q,K,V)=Concat(head1,...,head h )W O (12) Where: h is the number of attention heads, W O is the output linear transformation matrix; The feedforward neural network performs a nonlinear transformation on the features output by the self-attention mechanism. After the multi-head attention mechanism, the feature vector at each position passes through two layers of fully connected networks and adds a ReLU activation function: FFN(x)=W2(ReLU(W1x+b1))+b2 (14) Where: W1 and b1 are the weight and bias of the first fully connected network layer; W2 and b2 are the weights and biases of the second fully connected network layer; Residual connections and normalization Add residual connections and layer normalization after each sublayer: Output=LayerNorm(x+SubLayer(x)) (15) Where: LayerNorm(·) is the layer normalization calculation; SubLayer(x) is the output of a multi-head attention or feedforward neural network; The fully connected layer is specifically: The high-dimensional feature representation Z generated by the Transformer encoder L As the input of the fully connected layer, after one or more layers of linear transformation, the final output is the probability distribution of the rock burst danger level, and the Softmax activation function is used to output the level with the highest probability as the prediction result: p c =softmax(W d (x)+b d ) (16) Where: W d and b d are the weights and biases of the first fully connected layer; p c is the predicted probability that the rock burst category is predicted to be c.

5. The method for constructing a large-scale rock burst prediction model based on multimodal data according to claim 2, characterized in that: The details of step S3 are as follows: First, the rock burst hazard level (RL) is normalized to the [0, 1] interval. The mining information, geological structure, and prediction results data are also divided into the [0, 1] interval to facilitate the final assessment of the rock burst hazard. The mining information data is divided into criteria using the comprehensive index method, in which the specific criteria for classification of some factors are modified accordingly according to the actual situation; Mining information data impact factor W e The calculation is as follows: The geological structure data is divided into rules using the comprehensive index method to analyze the geological structure affected by the above geological data and mining data: Geological data: Mining data: Geological structure data impact factor W g Select the maximum value of geological data and mining comprehensive index, namely: W g =max{W g1 ,W g2 } (19) Deep learning data impact factor W m To output the danger level value, we first determine the danger level by the maximum probability value output by the model, and then divide the maximum probability value into 5 levels. At the same time, we divide the 4 danger level ranges into 5 levels corresponding to the maximum probability value range, which is used as the deep learning data influencing factor W. m ; Analysis shows that the maximum value of the probability of the model outputting the danger level (p c ) max The range is (0.25, 1], and according to the distribution characteristics of the model output probability, different deep learning data influence factors W are constructed. m ; The mining information data impact factor W obtained above is e , geological structure data influencing factor W g , Deep Learning Data Impact Factor W m , use weight division to determine the weight α of each part e , α g , α m , satisfying α e +α g +α m =1; First, considering that the weights should change dynamically with the mining process, we use the same time window as the precursor pattern sequence to count these three types of data and calculate the probability distribution of each type of data: Where: W k (l) represents the impact factor of mining information data, geological structure data, and deep learning data for the lth sample; b represents the total number of samples; Then, calculate the information entropy of each type of data: Where: E k Represents the information entropy of the k-th category data; ln(b) is the normalized coefficient of entropy; Therefore, the weight calculation formula for each part is: Where: α e , α g , α m The impact factor W of the information data e , geological structure data impact factor W g , Deep Learning Data Impact Factor W m The weight of Finally, the risk level RL for the predicted time period is calculated as: RL=a e W e +a g W g +a m W m (23).

Citation Information

Patent Citations

  • Assessment method for predicting underground rock burst danger of coal mine

    CN103244179A

  • Multi-parameter integrated monitoring and early-warning method for excavation working face

    CN105257339A