Coal mine earthquake time sequence feature prediction method based on multi-source data space-time diagram convolutional network

By using a combination method of multi-source data spatiotemporal graph convolution network and whale optimization algorithm in mineral seismic prediction, the problem of insufficient consideration of single data source dependence and complex physical characteristics in the existing technology is solved, and higher prediction accuracy and model performance are achieved.

CN120011792AActive Publication Date: 2025-05-16LIAONING UNIVERSITY

Patent Information

Application Number
CN202510081603.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing mineral seismic time series feature prediction models rely mostly on a single data source, ignoring the potential correlation and complementarity between different data sources, and the model is insufficient interpretability and failing to fully consider the complex physical characteristics of underground engineering, resulting in insufficient prediction accuracy.

Method used

A coal mine mine earthquake time series feature prediction method based on multi-source data spatiotemporal graph convolution network is designed. The features of micro-seismic and charge data are extracted through the feature fusion module. The spatiotemporal graph convolution module dynamically captures the correlation between time and space dimensions, and optimizes hyperparameters using whale optimization algorithm.

Benefits of technology

By fusion of multi-source data, dynamically capture spatiotemporal features and optimize hyperparameters, the accuracy of ore seismic prediction and model performance are significantly improved, and the understanding of changes in underground environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011792A_ABST
    Figure CN120011792A_ABST
Patent Text Reader

Abstract

The invention discloses a coal mine earthquake time sequence feature prediction method based on a multi-source data space-time diagram convolutional network, and belongs to the field of mine dynamic disaster prevention and control. The method comprises the following steps: firstly, fusing micro-seismic data and charge data, and extracting and integrating features of two data sources by using a feature fusion technology so as to more comprehensively capture dynamic features of an underground environment; then, constructing a graph structure, taking the sensors as nodes, and defining edges among the nodes by using Euclidean distances among the sensors so as to reflect spatial correlation of geological activities; and dynamically capturing the correlation between the time dimension and the space dimension through the space-time diagram convolutional network. And then, optimizing hyper-parameters in the model by using WOA to obtain an optimal parameter combination, and improving the performance and accuracy of the model. And finally, introducing a space-time attention mechanism, and dynamically adjusting data input by calculating a time attention matrix and a space attention matrix so as to better capture space-time characteristics and improve the accuracy of model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mine dynamic disaster prevention and control, and particularly relates to a prediction method for mine earthquake time series characteristics, and specifically relates to a coal mine earthquake time series characteristic prediction method based on a multi-source data spatiotemporal graph convolutional network. Background Art

[0002] Mine tremors are coal-rock dynamic phenomena that occur during mining and are one of the natural disasters in mines. How to reduce and mitigate accidents and disasters caused by mine tremors is an important research topic at present. One of the main ways to solve this problem is to monitor mine tremors in real time and continuously and predict the time series characteristics of mine tremor monitoring data. The rapid development and widespread application of machine learning methods have provided new ideas to make up for the shortcomings of numerical simulation methods. As a data-driven method, machine learning methods can effectively capture the nonlinear relationship between parameters in high-dimensional data and have been widely used in the field of geotechnical engineering.

[0003] In geological monitoring, different monitoring data sources usually provide multiple layers of information about underground structures and geological features. Various monitoring data such as microseismic, geoacoustic, stress, charge, and drill cuttings reflect the multifaceted characteristics of the underground environment, forming a multi-dimensional data system. In order to obtain more comprehensive and accurate information on underground environmental changes and thus improve the accuracy of prediction, it is necessary to make full use of these multi-source data for research. Mine earthquake prediction is equivalent to the problem of spatiotemporal data prediction. Data sensors record at fixed time points and fixed distribution locations. The observation results of adjacent nodes and timestamps are not independent, but dynamically related to each other. Therefore, effectively extracting the spatiotemporal correlation of data is the key to solving these problems. Although deep learning methods have shown potential in mine earthquake prediction, existing models mostly rely on a single data source, ignoring the potential correlation and complementarity between different data sources. In addition, the model is not interpretable enough and fails to fully consider the complex physical characteristics of underground engineering. The current theoretical model has the limitation of relying on a single indicator, and it is difficult to integrate laboratory and field monitoring status to achieve accurate prediction of the occurrence time of mine earthquakes. How to effectively extract the intrinsic patterns of nonlinear and complex spatiotemporal data and make accurate predictions is still a major challenge facing current research. Summary of the invention

[0004] In order to solve the problem of existing mine earthquake time series feature prediction, the present invention designs and implements a regression prediction method based on the whale optimization algorithm (WOA) to optimize the convolutional neural network (CNN) and the long short-term memory network (LSTM) combined model, and verifies the reliability of the method through practical application.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for predicting the characteristics of coal mine earthquake time series based on multi-source data spatiotemporal graph convolutional network, the steps of which are as follows:

[0007] Step 1 WM-STGCN network model construction: The overall model is divided into three parts: feature fusion, spatiotemporal graph convolution, and whale optimization algorithm hyperparameter optimization WOA;

[0008] In the feature fusion module, data features are extracted from the original microseismic data and charge data through a feature encoder, and single-element feature extraction is performed to obtain the global time series features of each data type. The relationship between the data features of the two is learned to obtain the local interaction features between microseismic and charge. The global time series features of each data type are fused with the local interaction features between microseismic and charge to obtain the fusion features, thus realizing the data fusion of microseismic and charge.

[0009] The spatiotemporal graph convolution module dynamically captures the correlation between the time and space dimensions by calculating the time and space attention matrices, and inputs the data adjusted by the attention mechanism into the spatiotemporal convolution layer to obtain dependencies from different dimensions;

[0010] The WOA hyperparameter optimization module optimizes the hyperparameters required in the model and finds the optimal parameter combination.

[0011] The overall structure and usage of the WM-STGCN model are as follows:

[0012] First, the time series data is cut into three time dimension sequence segments along the time axis according to the length of hour, day, and week, which are used as the input of hour component, day component, and week component respectively;

[0013] Secondly, the features of microseismic data and charge data are extracted from the sensor data through the feature encoder. The global time series features in the single field data and the local interaction features between the multi-field data are learned through the intra-field data relationship learning module and the inter-field data relationship learning module. Finally, the fusion of intra-field features and multi-field interaction features is realized through the multi-feature fusion module.

[0014] Then the model is composed of multiple spatiotemporal blocks, each of which has a spatiotemporal attention module and a spatiotemporal convolution module, and a residual learning framework is adopted in each component;

[0015] In the hyperparameter optimization part, the whale optimization algorithm is used to optimize the learning rate, number of spatiotemporal convolution kernels, number of model blocks, and convolution order hyperparameters for data fusion and spatiotemporal graph convolution to find the best parameter combination;

[0016] Finally, the outputs of the three components of hour, day, and week are further merged according to the parameter matrix to obtain the final prediction result.

[0017] Step 2 Data processing:

[0018] Set the maximum sampling frequency of the sensor to q times per day, the current time to 0, and the prediction window size to T p , the intercept length along the time axis is T h , T d and T w The three sequence fragments are used as the input of the three periodic components of time, day and week, respectively. h , T d and T w All T p multiples of.

[0019] In the step 2, the specific method is:

[0020] Set the maximum sampling frequency of the sensor to q times per day, the current time to 0, and the prediction window size to T p . The length of the intercept along the time axis is T h , T d and T w The three sequence fragments are used as the input of the three periodic components of time, day and week, respectively. h , T d and T w All T p multiples of.

[0021] 2.1) Time period sequence

[0022]

[0023] Where: T h is a time-period sequence fragment, t0 represents the current time, N and F are the dimension sizes, representing the number of samples and the number of features respectively;

[0024] As shown in formula (1), the time period series is a historical time series directly adjacent to the prediction period. For mine earthquake data, the occurrence and end of mine earthquake events are gradual. Therefore, the newly generated mine earthquake data will have an impact on future data information.

[0025] 2.2) Daily cycle sequence

[0026]

[0027] Where: T d is a sequence segment of a daily cycle, q is the maximum number of sampling times of the sensor per day, T p is the predicted window size, t0, N, F are the same as above;

[0028] As shown in formula (2), the daily cycle sequence is composed of the data of the past few days in the same time period as the prediction period. Considering the daily production of the mining area, the mine earthquake data will change periodically with days;

[0029] 2.3) Weekly cycle sequence

[0030]

[0031] Where: T w is a sequence segment of a daily cycle, q, T p , t0, N, F are the same as above;

[0032] As shown in formula (3), the weekly cycle sequence is composed of segments of recent weeks. The coal mine working pattern on Monday is similar to the coal mine working pattern on Mondays in history, but is quite different from the coal mine working pattern on weekends.

[0033] Step 3: Feature fusion: Multi-feature fusion of microseismic data and charge data is performed to process data of different modes within the same framework.

[0034] 3.1) Feature extraction

[0035] A data feature extraction method based on deep learning is used, combined with a bidirectional long short-term memory network and an average pooling method to capture timing information and balance the output at different times. A BiLSTM network is used to capture the timing relationship in sensor data. Microseismic and charge data are processed by independent networks respectively. The BiLSTM network has hidden states in both the forward and backward directions. For microseismic data, after data processing in step 2, X h , X d , X w , here we use micor t (i) represents, as input data, and then obtained at each sensor position i, the representation is shown in formula (4)(5)(6):

[0036]

[0037] Among them, micor t (i) represents the input of microseismic data, and are the forward and backward hidden states, respectively. is the hidden state at the corresponding moment; in order to balance the output at different moments, the feature of each moment t Perform average pooling, as shown in formula (7):

[0038]

[0039] Where T represents the total number of moments of the time series data. It is the feature representation after average pooling;

[0040] Perform the above processing on the charge data to complete feature extraction;

[0041] 3.2) Single Data Relationship Learning

[0042] By constructing a fully connected graph attention network, we extract internal global time series features for microseismic and charge data respectively. We divide the features of microseismic and charge data to form their own fully connected graphs. We use the self-attention mechanism of the Transformer network to divide each data feature into several nodes to ensure that each node has an adjacency relationship with other nodes in the field. We perform multi-head neighbor aggregation on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in Equations (8), (9), and (10):

[0043]

[0044] Among them, Q represents the weight of the edge, which is the self-attention coefficient value calculated by the self-attention mechanism in the Transformer network; X μ Represents the characteristics before aggregation; represents the aggregated features, i = 1,...,h, h represents the number of self-attention heads, μ∈{m,c}, m,c represent microseisms and charges respectively; d k express or The dimension of , softmax represents the activation function, Both represent weight matrices, and || represents cascade operations. Since the fully connected graph attention network is a multi-layer graph learning network, multi-layer graph learning is required to obtain the final global temporal characteristics of single-factor data.

[0045] 3.3) Global Data Relationship Learning

[0046] In order to capture the complex relationship between microseismic and electric charge, a sparse graph attention network is constructed to learn the data relationship of the data features of microseismic and electric charge. The data features of microseismic and electric charge are divided into several nodes by using the sparse graph attention network. For each node, the adjacency relationship with the nodes in another data type is calculated. Each node retains X attention coefficient values ​​from large to small, and the remaining attention coefficient values ​​are set to zero to form a sparse graph in the field. Multi-head neighbor aggregation is performed on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in formulas (11), (12), and (13):

[0047]

[0048] Among them, μ m ,μ c represent microseism and electric charge respectively; It represents the characteristics of the nodes in the microseism after aggregating the nodes in the charge that have adjacent relationships with them; and represents the features before aggregation, Represents the weight matrix, performs multi-layer graph learning to obtain the final local interaction features between data;

[0049] The global time series features of the single factor data are fused with the local interaction features between the data to obtain the fusion features, realize the fusion of microseismic charge data, and finally the fusion feature X is shown in formula (14):

[0050]

[0051] Step 4: Hyperparameter optimization

[0052] For the model constructed in step 1, an optimization algorithm needs to be selected. The whale optimization algorithm is selected as a tool for hyperparameter optimization. Through the intelligent search mechanism, it evolves in the parameter space to find the optimal hyperparameter combination.

[0053] In the step 4, the specific method is:

[0054] Taking the initial data as the input of the model in step 1, the specific formula of the algorithm is as follows:

[0055] X i (t+1)=X i (t)+A·rand()·dist(X i (t),

[0056] X best (t))-B·dist(X i (t),X rand (t)) (15)

[0057] Among them, X i (t) represents the position of the ith whale at time t, A and B are the control parameters of the algorithm, is the random number generation function, dist(·,·) represents the distance function between two points, and X best (t) and X rand (t) are the global optimal position and random position at the current moment respectively.

[0058] Step 5: Spatiotemporal graph convolution

[0059] After obtaining the optimal hyperparameter combination, the spatial and temporal attention mechanism is introduced, and then the spatiotemporal graph convolution calculation is performed on the fused feature value X obtained in step 2 in combination with the hyperparameters in step 4.

[0060] 5.1) Spatial Attention Mechanism

[0061] In the spatial dimension, the influence of different channels on each other is dynamic. By calculating the spatial attention matrix, the attention mechanism is used to dynamically capture this correlation, as shown in equations (16) and (17):

[0062]

[0063] in, is the input of the rth spatiotemporal block, C (r-1) is the number of channels of input data, b s ,V s ∈R NxN , is a learnable parameter, σ is an activation function;

[0064] 5.2) Temporal Attention Mechanism

[0065] In the time dimension, the data on different time slices of the same channel also have dynamic correlations. The attention mechanism is used to dynamically capture this correlation, as shown in formulas (18) and (19):

[0066]

[0067] in, U1∈R N , is a learnable parameter, σ is an activation function;

[0068] 5.3) Spatial Dependency Modeling

[0069] Microseismic data and charge data are essentially graph structures. The features of each node are regarded as signals on the graph. Graph convolution based on graph spectrum theory is used to process each time slice. The signal correlation of mine seismic data in the spatial dimension is used to convert the graph into an algebraic form and analyze the topological properties of the graph. The convolution operation of the graph signal is equal to the product of these signals transformed into the spectral domain by the graph Fourier transform. When the scale of the graph is large, the Chebnet model is used, as shown in formula (20). After adding the spatial attention mechanism, it is shown in formula (21):

[0070]

[0071] 5.4) Time-dependent modeling

[0072] After the graph convolution operation captures the adjacent information of each node on the graph in the spatial dimension, it further stacks the standard convolution layer in the time dimension to update the node signal by merging the information on adjacent time slices. After capturing the neighborhood information in the spatial dimension, a standard convolution layer is added in the temporal dimension to update the node information by merging the domain time slice information, as shown in formula (22):

[0073]

[0074] Step 6: Model training and prediction

[0075] Train the model using the training set and adjust the model parameters to minimize the prediction error;

[0076] The input is the training set and the defined graph node relationship matrix. First, the learning parameters of the model are initialized, including the definition of the graph convolution layer, attention mechanism, and fully connected layer. Then, batch data are selected for training through iteration, including microseismic data and charge data. During the iteration process, multi-feature fusion is performed on the data, hyperparameter optimization is performed using the WOA algorithm, and spatiotemporal dependencies are obtained through the spatiotemporal attention mechanism. Finally, the prediction data sequence is output through the fully connected layer.

[0077] The beneficial effects created by the present invention are as follows: the present invention fuses microseismic data and charge data, and uses feature fusion technology to extract and integrate the features of the two data sources to more comprehensively capture the dynamic characteristics of the underground environment. Then, a graph structure is constructed, with sensors as nodes, and the Euclidean distance between sensors is used to define the edges between nodes to reflect the spatial correlation of geological activities. The correlation between the time and space dimensions is dynamically captured through the spatiotemporal graph convolutional network. Next, the whale optimization algorithm WOA is used to optimize the hyperparameters in the model to find the optimal parameter combination and improve the performance and accuracy of the model. Finally, the spatiotemporal attention mechanism is introduced to dynamically adjust the data input by calculating the time and space attention matrix to better capture the spatiotemporal characteristics and improve the accuracy of the model prediction. The high precision and good applicability of the model in mine earthquake prediction are verified by experiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 : Experimental system diagram of the present invention;

[0079] Figure 2 : A line graph comparing the predicted value and the true value of the present invention;

[0080] Figure 3 : Comparison of model training time cost;

[0081] Figure 4 : The traffic flow prediction result comparison line chart on PEMS04 of the present invention;

[0082] Figure 5 :3 -1 105 working surface plan position diagram;

[0083] Figure 6 : The mine earthquake event early warning map of the present invention;

[0084] Figure 7 : Step 6 training process diagram of the present invention. DETAILED DESCRIPTION

[0085] The technical solutions in the embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. The following embodiments are applicable to the present invention, but are not intended to limit the scope of the present invention.

[0086] A method for predicting the characteristics of coal mine earthquake time series based on multi-source data spatiotemporal graph convolutional network, the steps of which are as follows:

[0087] Step 1 WM-STGCN network model construction

[0088] The proposed WM-STGCN overall framework can be divided into three parts: feature fusion, spatiotemporal graph convolution, and Whale Optimization Algorithm (WOA) hyperparameter optimization. In the feature fusion module, the feature encoder is used to extract data features from the original microseismic data and charge data, and single-element feature extraction is performed to obtain the global time series features of each data type. The relationship between the data features of the two is learned to obtain the local interaction features between microseismic and charge. The global time series features of each data type are fused with the local interaction features between microseismic and charge to obtain the fusion features, thereby realizing the data fusion of microseismic and charge. The spatiotemporal graph convolution module dynamically captures the correlation between the two dimensions by calculating the time and space attention matrix, and inputs the data adjusted by the attention mechanism into the spatiotemporal convolution layer to obtain dependencies from different dimensions. The WOA hyperparameter optimization module performs parameter optimization on the hyperparameters required in the model and finds the optimal parameter combination, thereby improving the performance and accuracy of the model.

[0089] The overall structure and usage of the WM-STGCN model can be summarized as follows:

[0090] First, the time series data is truncated into three time dimension sequence segments along the time axis according to the length of hours, days, and weeks, which are used as the input of hour component, day component, and week component respectively.

[0091] Secondly, the features of microseismic data and charge data are extracted from the sensor data through the feature encoder. The global time series features in the single field data and the local interaction features between the multi-field data are learned through the intra-field data relationship learning module and the inter-field data relationship learning module. Finally, the intra-field features and multi-field interaction features are fused through the multi-feature fusion module.

[0092] The model is then composed of multiple spatiotemporal blocks. In each spatiotemporal block, there is a spatiotemporal attention module and a spatiotemporal convolution module. In order to optimize the training efficiency, a residual learning framework is adopted in each component.

[0093] In the hyperparameter optimization part, we optimize the model hyperparameters required for data fusion and spatiotemporal graph convolution to find the best parameter combination to further improve the model performance. Specifically, we use the whale optimization algorithm to optimize the learning rate, number of spatiotemporal convolution kernels, number of model blocks, convolution order and other hyperparameters.

[0094] Finally, the outputs of the three components are further combined according to the parameter matrix to obtain the final prediction result.

[0095] Step 2 Data processing

[0096] Set the maximum sampling frequency of the sensor to q times per day, the current time to 0, and the prediction window size to T p . The length of the intercept along the time axis is T h , T d and T w The three sequence fragments are used as the input of the three periodic components of time, day and week, respectively. h , T d and T w All T p multiples of.

[0097] (1) Time period sequence

[0098]

[0099] As shown in formula (1), the time period series is a historical time series directly adjacent to the prediction period. For mine earthquake data, because of its strong autocorrelation, the occurrence and end of mine earthquake events are gradual. Therefore, the newly generated mine earthquake data will inevitably affect future data information.

[0100] (2) Daily cycle sequence

[0101]

[0102] As shown in formula (2), the daily cycle sequence is composed of the data of the past few days in the same time period as the prediction period. Considering the daily production of the mining area, the mine earthquake data may change periodically on a daily basis, such as starting work, leaving work, and resting on time every day.

[0103] (3) Weekly cycle sequence

[0104]

[0105] As shown in formula (3), the weekly cycle sequence consists of segments of recent weeks. Generally speaking, the coal mine working pattern on Monday has certain similarities with the coal mine working pattern on Mondays in history, but may be quite different from the coal mine working pattern on weekends.

[0106] Step 3: Feature Fusion

[0107] Microseismic data and charge data provide complementary information in the geological field, but due to the heterogeneity of their respective characteristics, how to effectively integrate them becomes a key issue. Through multi-feature fusion, data of different modalities can be processed within the same framework, improving the overall coherence and comprehensiveness of the information.

[0108] (1) Feature extraction

[0109] This paper adopts a data feature extraction method based on deep learning, aiming to extract meaningful data features of microseismicity and charge from raw sensor data. This method combines the bidirectional long short-term memory network (BiLSTM) and average pooling technology to capture timing information and balance the output at different times. In order to effectively capture the temporal relationship in sensor data, BiLSTM networks are applied respectively, in which microseismicity and charge data are processed by independent networks respectively. The BiLSTM network has hidden states in both forward and backward directions. For microseismicity data, after data processing in step 2, X is obtained. h , X d , X w , here we use micor t (i) is represented as input data. Then, at each sensor position i, the representation is shown in formula (4), (5), and (6).

[0110]

[0111] Among them, micor t (i) represents the input of microseismic data, and are the forward and backward hidden states, respectively. is the hidden state at the corresponding moment. In order to balance the output at different moments, it is necessary to Perform average pooling as shown in formula (7).

[0112]

[0113] Where T represents the total number of moments of the time series data. It is the feature representation after average pooling. Similar processing is performed on the charge data to complete feature extraction.

[0114] (2) Single Data Relationship Learning

[0115] In order to deeply explore the internal correlation of microseismic and charge data features, this paper conducts internal data relationship learning and realizes internal global time series feature extraction for microseismic and charge data by constructing a fully connected graph attention network. The microseismic and charge data features are divided to form their own fully connected graphs. The self-attention mechanism of the Transformer network is used to divide each data feature into several nodes to ensure that each node has an adjacency relationship with other nodes in the field. Multi-head neighbor aggregation is performed on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in Equations (8), (9), and (10).

[0116]

[0117]

[0118] Among them, Q represents the weight of the edge, which is the self-attention coefficient value calculated by the self-attention mechanism in the Transformer network; X μ Represents the characteristics before aggregation; represents the aggregated features, i = 1,...,h, h represents the number of self-attention heads, μ∈{m,c}, m,c represent microseisms and charges respectively; d k express or The dimension of , softmax represents the activation function, Both represent weight matrices, and || represents cascade operations. Since the fully connected graph attention network is a multi-layer graph learning network, multi-layer graph learning is required to obtain the final global temporal features of single-factor data.

[0119] (3) Global data relationship learning

[0120] In order to capture the complex relationship between microseisms and charges, a sparse graph attention network is constructed to learn the data relationship of the data features of microseisms and charges. Using the sparse graph attention network, the data features of microseisms and charges are divided into several nodes respectively. For each node, its adjacency relationship with the nodes in another data type is calculated. Specifically, each node retains X attention coefficient values ​​from large to small, and the remaining attention coefficient values ​​are set to zero to form a sparse graph in the field. Multi-head neighbor aggregation is performed on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in formulas (11)(12)(13).

[0121]

[0122] Among them, μ m ,μ c represent microseism and electric charge respectively; It represents the characteristics of the nodes in the microseism after aggregating the nodes in the charge that have adjacent relationships with them; and represents the features before aggregation, Represents the weight matrix. Perform multi-layer graph learning to obtain the final local interaction features between data.

[0123] The global time series features of single factor data are fused with the local interaction features between data to obtain the fusion features, thus realizing the fusion of microseismic charge data. The final fusion feature X is shown in formula (14).

[0124]

[0125] Step 4: Hyperparameter optimization

[0126] The performance of deep learning models is often affected by hyperparameters, and the selection of hyperparameters is crucial to the training effect of the model. Therefore, hyperparameter optimization has become one of the key paths to optimize deep learning models. This process aims to make the model better adapt to the data and improve the training effect and generalization ability by adjusting key hyperparameters such as learning rate, number of convolution kernels, and number of model blocks. A reasonably selected hyperparameter combination can accelerate the convergence speed of the model, improve the generalization ability of the model, and reduce the risk of overfitting or underfitting. If a fixed hyperparameter combination is used without optimization, the model may fall into a local optimal solution during the training process and fail to fully explore the potential laws of the data. This may lead to performance degradation, insufficient generalization ability, and even poor performance on the test set. The lack of hyperparameter optimization may prevent the model from fully exerting its advantages, so in deep learning tasks, the optimization of hyperparameters is particularly critical.

[0127] In order to deal with the problems caused by hyperparameters, the model constructed in step 1 needs to select an optimization algorithm. The present invention selects the whale optimization algorithm as a tool for hyperparameter optimization. The whale optimization algorithm has the ability of global search by simulating the foraging behavior of whale groups, and can efficiently find the optimal solution in the multi-dimensional parameter space. This algorithm draws on the collective behavior of whale groups in nature, and evolves in the parameter space through an intelligent search mechanism in order to find the optimal combination of hyperparameters. The core idea of ​​the whale optimization algorithm is to simulate the foraging process of whale groups and continuously adjust the parameters to approach the optimal solution. Its mathematical expression includes an intelligent search mechanism for parameter space, and realizes efficient hyperparameter optimization through randomness and tracking of the global optimal solution. Here, the initial data is used as the input of the model in step 1, and the specific formula of the algorithm is as follows:

[0128] X i (t+1)=X i (t)+A·rand()·dist(X i (t),

[0129] X best (t))-B·dist(X i (t),X rand (t)) (15)

[0130] Among them, X i (t) represents the position of the ith whale at time t, A and B are the control parameters of the algorithm, is the random number generation function, dist(·,·) represents the distance function between two points, and X best (t) and X rand (t) are the global optimal position and random position at the current moment respectively.

[0131] The reason for choosing the whale optimization algorithm is its advantage of global search. Compared with other optimization algorithms, the whale optimization algorithm has a more global search feature, which enables it to efficiently find the optimal solution in a wide range of parameter spaces. The global search nature of the whale optimization algorithm enables it to find better parameter combinations in complex deep learning tasks and improve the performance of the model. This global search capability plays a key role in the diversity and uncertainty of the hyperparameter space, helping the model better adapt to different data distributions and task scenarios.

[0132] Step 5: Spatiotemporal graph convolution

[0133] After obtaining the optimal hyperparameter combination, the spatial and temporal attention mechanism is introduced, and then the spatiotemporal graph convolution calculation is performed on the fused feature value X obtained in step 2 in combination with the hyperparameters in step 4.

[0134] (1) Spatial Attention Mechanism

[0135] In the spatial dimension, the influence of different channels on each other is dynamic, so the attention mechanism is used to dynamically capture this correlation, which is achieved by calculating the spatial attention matrix, as shown in Equations (16) and (17).

[0136]

[0137] in, is the input of the rth spatiotemporal block, C (r-1) is the number of channels of input data, b s ,V s ∈R NxN , is a learnable parameter and σ is an activation function.

[0138] (2) Temporal Attention Mechanism

[0139] In the time dimension, there is also dynamic correlation in the data of the same channel at different time slices. This scheme uses the attention mechanism to dynamically capture this correlation, as shown in equations (18) and (19).

[0140]

[0141] in, U1∈R N , is a learnable parameter and σ is an activation function.

[0142] (3) Spatial Dependency Modeling

[0143] Microseismic data and charge data are essentially graph structures. The features of each node can be regarded as signals on the graph. Therefore, in order to make full use of such topological characteristics, it is necessary to use graph convolution based on graph spectrum theory on each time slice to process the signal correlation of mine seismic data in the spatial dimension. The spectral method converts the graph into an algebraic form and analyzes the topological properties of the graph. Since the convolution operation of the graph signal is equal to the product of these signals transformed into the spectral domain by the graph Fourier transform, however, when the scale of the graph is large, the direct eigenvalue decomposition of the Laplace matrix is ​​very expensive. Therefore, this paper uses the Chebnet model to solve this problem, as shown in Equation (20). After adding the spatial attention mechanism, it is shown in Equation (21).

[0144]

[0145] (4) Time-dependent modeling

[0146] After the graph convolution operation captures the adjacent information of each node on the graph in the spatial dimension, a standard convolution layer in the time dimension is further stacked to update the node signal by merging the information on adjacent time slices. After capturing the neighborhood information in the spatial dimension, a standard convolution layer is added in the temporal dimension to update the node information by merging the domain time slice information, as shown in formula (22).

[0147]

[0148] Step 6: Model training and prediction process

[0149] The model is trained through the training set, and the model parameters are continuously adjusted to minimize the prediction error. The input is the training set and the defined graph node relationship matrix. The algorithm first initializes the learning parameters of the model, including defining the graph convolution layer, attention mechanism, etc. Then, batch data is selected for training through iteration, including microseismic data and charge data. During the iteration process, multi-feature fusion is performed on the data, hyperparameter optimization is performed using the WOA algorithm, and spatiotemporal dependencies are obtained through the spatiotemporal attention mechanism. Finally, the prediction data sequence is output through the fully connected layer. The training loss is calculated based on the objective function, and the model parameters are updated through back propagation. The running time of the training algorithm is related to the iteration cycle, the size of the training data, and the model parameters. The overall running complexity is O(n 2 ), the model training process is as follows Figure 7 shown.

[0150] Example 1

[0151] In order to verify the present invention, an experiment was conducted, and a mine earthquake dataset HYSK collected from a mine in Liaoning from April 2022 to July 2022 was selected. The dataset contains three-component waveform microseismic data of 100 mine earthquake events from 9 microseismic stations and charge data from 6 charge probes. The microseismic data acquisition frequency is 5KHz, and the charge data is 1KHz. In order to optimize the data quality, processing steps such as zero padding, long and short time window method to align P waves, downsampling and Z-Score normalization were used to ensure the integrity and consistency of the data. Finally, the dataset was divided into training set and test set in a ratio of 7:3, and input into the training model. The model framework is as follows Figure 1 shown.

[0152] After the data is input into the model, feature fusion is first performed. The microseismic and charge data features are extracted through deep learning methods, and the internal and global relationships between the data are learned using the fully connected graph attention network and the sparse graph attention network to achieve feature fusion. The whale optimization algorithm is further used to optimize the hyperparameters to find the optimal parameter combination. The spatial and temporal attention mechanism and graph convolution operation are used to capture the spatiotemporal dependencies of the data, and finally the prediction results are output, such as Figure 2 shown.

[0153] The mean squared error (MSE) was selected as the model objective function, and the back propagation algorithm was used for training. WM-TGCN was designed as the input sequence length to achieve multi-step prediction. The Adam Optimezer optimizer was used to reduce the error. The mean absolute error (MAE), mean absolute percentage error (MAPE) and root mean square error (RMSE) were used to measure the performance of different models.

[0154] Evaluate the model training time and cost to verify the practical applicability of the model. Figure 3As shown in the figure, in the comparison of model time cost, WM-STGCN may show a slightly higher time cost than other models. However, it is worth noting that although the time cost of WM-STGCN has increased, its training time has not increased exponentially, but has been moderately improved within an acceptable range. This feature makes WM-STGCN still maintain good feasibility and practicality in practical applications.

[0155] Finally, a generalization experiment was conducted to verify the effectiveness of the model in other fields. The public data set PEMS04 in the field of transportation was selected. This data set records multi-dimensional data measured by sensors on California highways from January 1, 2018 to February 28, 2018, including total flow, average speed, and average occupancy. The control model uses the STGCN model. In WM-STGCN, the total flow is predicted by combining data such as average speed. The experimental results are as follows: Figure 4 As shown in the figure, compared with STGCN, WM-STGCN achieves the best learning effect. MAE and RMSE are both optimal in the prediction evaluation results. As the number of prediction steps increases, the error of the WM-STGCN model does not explode, which shows that it has strong robustness and proves that the proposed model has strong generalization ability.

[0156] An example of the present invention is given below in conjunction with the accompanying drawings:

[0157] (1) In order to verify the effectiveness of the WM-STGCN model in actual mine earthquake prediction, it was actually applied to the 3-1105 working face of a coal mine for mine earthquake prediction and early warning. The 3-1105 working face is located in the middle of the 3-11 mining area and the southwest of the mine industrial square. The terrain is a slope with high northeast and low southwest. The highest point is located near the surface corresponding to the middle of the 18-12 and 19-10 boreholes, with an altitude of 1400.8m. The lowest point is located on the ground corresponding to the main withdrawal channel of the working face, with an altitude of 1348.5m. The maximum elevation difference of the terrain location is 52.3m. To the southeast of the 3-1105 working face is the auxiliary transportation drift (unexcavated) of the 3-1107 working face, and to the northwest is a 6-meter coal pillar separated from the 3-1103 goaf (966 meters along the goaf strike length), part of which is solid coal, and the other surrounding directions are solid coal. The plane position diagram of the 3-1105 working face is shown in the figure below. Figure 5 shown.

[0158] (2) The monitoring data of the working face in June 2022 were selected for time series feature prediction, and divided into four level intervals of 1, 2, 3, and 4 according to the maximum energy value of the microseismic on that day, corresponding to four impact levels respectively.

[0159] (3) Finally, we get Figure 6 The early warning map shown shows the comparison between the actual hazard level and the predicted hazard level of the 3-1105 working face during June 2022. Through detailed analysis, it can be clearly seen that the energy-derived hazard level calculated based on the actual data is highly consistent with the energy-derived hazard level calculated by the predicted data. Especially when the predicted hazard level is high, the prediction accuracy of actual mine earthquake events is as high as 80%. At the same time, on June 8, although the actual hazard level was rated as medium, the data results of the prediction model were high hazard levels, and a mine earthquake event did occur subsequently. This example not only verifies the accuracy of the WM-STGCN model proposed in this paper in predicting mine earthquake events, but also demonstrates its effectiveness and reliability in dealing with potential risks.

Claims

1. A method for predicting the time series characteristics of coal mine earthquakes based on a multi-source data spatiotemporal graph convolutional network, characterized by: Step 1 WM-STGCN network model construction: The overall model is divided into three parts: feature fusion, spatiotemporal graph convolution, and whale optimization algorithm hyperparameter optimization WOA; In the feature fusion module, data features are extracted from the original microseismic data and charge data through a feature encoder, and single-element feature extraction is performed to obtain the global time series features of each data type. The relationship between the data features of the two is learned to obtain the local interaction features between microseismic and charge. The global time series features of each data type are fused with the local interaction features between microseismic and charge to obtain the fusion features, thus realizing the data fusion of microseismic and charge. The spatiotemporal graph convolution module dynamically captures the correlation between the time and space dimensions by calculating the time and space attention matrices, and inputs the data adjusted by the attention mechanism into the spatiotemporal convolution layer to obtain dependencies from different dimensions; The WOA hyperparameter optimization module optimizes the hyperparameters required in the model and finds the optimal parameter combination; Step 2 Data processing: Set the maximum sampling frequency of the sensor to q times per day, the current time to 0, and the prediction window size to T p , the intercept length along the time axis is T h , T d and T w The three sequence fragments are used as the input of the three periodic components of time, day and week, respectively. h , T d and T w All T p multiples of; Step 3: Feature fusion: Multi-feature fusion of microseismic data and charge data, processing data of different modes in the same framework; Step 4: Hyperparameter optimization For the model constructed in step 1, an optimization algorithm needs to be selected. The whale optimization algorithm is selected as a tool for hyperparameter optimization. The optimal hyperparameter combination is found by evolving in the parameter space through an intelligent search mechanism. Step 5: Spatiotemporal graph convolution After obtaining the optimal hyperparameter combination, the spatial and temporal attention mechanism is introduced, and then the spatiotemporal graph convolution calculation is performed on the fused feature value X obtained in step 2 in combination with the hyperparameters in step 4; Step 6: Model training and prediction Train the model using the training set and adjust the model parameters to minimize the prediction error; The input is the training set and the defined graph node relationship matrix. First, the learning parameters of the model are initialized, including the definition of the graph convolution layer, attention mechanism, and fully connected layer. Then, batch data are iteratively selected for training, including microseismic data and charge data; In the iterative process, the data is fused with multiple features, the WOA algorithm is used to optimize the hyperparameters, and the spatiotemporal dependency is obtained through the spatiotemporal attention mechanism. Finally, the prediction data sequence is output through the fully connected layer.

2. The method for predicting coal mine earthquake time series characteristics based on multi-source data spatiotemporal graph convolutional network according to claim 1 is characterized by: The overall structure and usage of the WM-STGCN model are as follows: First, the time series data is cut into three time dimension sequence segments along the time axis according to the length of hour, day, and week, which are used as the input of hour component, day component, and week component respectively; Secondly, the features of microseismic data and charge data are extracted from the sensor data through the feature encoder. The global time series features in the single field data and the local interaction features between the multi-field data are learned through the intra-field data relationship learning module and the inter-field data relationship learning module. Finally, the fusion of intra-field features and multi-field interaction features is realized through the multi-feature fusion module. Then the model is composed of multiple spatiotemporal blocks, each of which has a spatiotemporal attention module and a spatiotemporal convolution module, and a residual learning framework is adopted in each component; In the hyperparameter optimization part, the whale optimization algorithm is used to optimize the learning rate, number of spatiotemporal convolution kernels, number of model blocks, and convolution order hyperparameters for data fusion and spatiotemporal graph convolution to find the best parameter combination; Finally, the outputs of the three components of hour, day, and week are further merged according to the parameter matrix to obtain the final prediction result.

3. The method for predicting coal mine earthquake time series characteristics based on multi-source data spatiotemporal graph convolutional network according to claim 1 is characterized by: In the step 2, the specific method is: Set the maximum sampling frequency of the sensor to q times per day, the current time to 0, and the prediction window size to T p ; The intercept length along the time axis is T h , T d and T w The three sequence fragments are used as the input of the three periodic components of time, day and week, respectively. h , T d and T w All T p multiples of; 2.1) Time period sequence Where: T h is a time-period sequence fragment, t0 represents the current time, N and F are the dimension sizes, representing the number of samples and the number of features respectively; As shown in formula (1), the time period series is a historical time series directly adjacent to the prediction period. For mine earthquake data, the occurrence and end of mine earthquake events are gradual. Therefore, the newly generated mine earthquake data will have an impact on future data information. 2.2) Daily cycle sequence Where: T d is a sequence segment of a daily cycle, q is the maximum number of sampling times of the sensor per day, T p is the predicted window size, t0, N, F are the same as above; As shown in formula (2), the daily cycle sequence is composed of the data of the past few days in the same time period as the prediction period. Considering the daily production of the mining area, the mine earthquake data will change periodically with days; 2.3) Weekly cycle sequence Where: T w is a sequence segment of a daily cycle, q, T p , t0, N, F are the same as above; As shown in formula (3), the weekly cycle sequence is composed of segments of recent weeks. The coal mine working pattern on Monday is similar to the coal mine working pattern on Mondays in history, but is quite different from the coal mine working pattern on weekends.

4. The method for predicting coal mine earthquake time series characteristics based on multi-source data spatiotemporal graph convolutional network according to claim 1 is characterized in that: The specific method in step 3 is: 3.1) Feature extraction A data feature extraction method based on deep learning is used, combined with a bidirectional long short-term memory network and an average pooling method to capture timing information and balance the output at different times. A BiLSTM network is used to capture the timing relationship in sensor data. Microseismic and charge data are processed by independent networks respectively. The BiLSTM network has hidden states in both the forward and backward directions. For microseismic data, after data processing in step 2, X h , X d , X w , here we use micor t (i) represents, as input data, and then obtained at each sensor position i, the representation is shown in formula (4)(5)(6): Among them, micor t (i) represents the input of microseismic data, and are the forward and backward hidden states, respectively. is the hidden state at the corresponding moment; in order to balance the output at different moments, the feature of each moment t Perform average pooling, as shown in formula (7): Where T represents the total number of moments of the time series data. It is the feature representation after average pooling; Perform the above processing on the charge data to complete feature extraction; 3.2) Single Data Relationship Learning By constructing a fully connected graph attention network, we extract internal global time series features for microseismic and charge data respectively. We divide the features of microseismic and charge data to form their own fully connected graphs. We use the self-attention mechanism of the Transformer network to divide each data feature into several nodes to ensure that each node has an adjacency relationship with other nodes in the field. We perform multi-head neighbor aggregation on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in Equations (8), (9), and (10): Among them, Q represents the weight of the edge, which is the self-attention coefficient value calculated by the self-attention mechanism in the Transformer network; X μ Represents the characteristics before aggregation; represents the aggregated features, i = 1,...,h, h represents the number of self-attention heads, μ∈{m,c}, m,c represent microseisms and charges respectively; d k express or The dimension of , softmax represents the activation function, Both represent weight matrices, and || represents cascade operations. Since the fully connected graph attention network is a multi-layer graph learning network, multi-layer graph learning is required to obtain the final global temporal characteristics of single-factor data. 3.3) Global Data Relationship Learning In order to capture the complex relationship between microseismic and electric charge, a sparse graph attention network is constructed to learn the data relationship of the data features of microseismic and electric charge. The data features of microseismic and electric charge are divided into several nodes by using the sparse graph attention network. For each node, the adjacency relationship with the nodes in another data type is calculated. Each node retains X attention coefficient values ​​from large to small, and the remaining attention coefficient values ​​are set to zero to form a sparse graph in the field. Multi-head neighbor aggregation is performed on each node and all adjacent nodes of each fully connected graph to obtain the aggregated features, as shown in formulas (11), (12), and (13): Among them, μ m ,μ c represent microseism and electric charge respectively; It represents the characteristics of the nodes in the microseism after aggregating the nodes in the charge that have adjacent relationships with them; and represents the features before aggregation, Represents the weight matrix, performs multi-layer graph learning to obtain the final local interaction features between data; The global time series features of the single factor data are fused with the local interaction features between the data to obtain the fusion features, realize the fusion of microseismic charge data, and finally the fusion feature X is shown in formula (14):

5. The method for predicting coal mine earthquake time series characteristics based on multi-source data spatiotemporal graph convolutional network according to claim 1 is characterized by: In the step 4, the specific method is: Taking the initial data as the input of the model in step 1, the specific formula of the algorithm is as follows: X i (t+1)=X i (t)+A·rand()·dist(X i (t), X best (t))-B·dist(X i (t),X rand (t)) (15) Among them, X i (t) represents the position of the ith whale at time t, A and B are the control parameters of the algorithm, is the random number generation function, dist(·,·) represents the distance function between two points, and X best (t) and X rand (t) are the global optimal position and random position at the current moment respectively.

6. The method for predicting coal mine earthquake time series characteristics based on multi-source data spatiotemporal graph convolutional network according to claim 1 is characterized by: In the step 5, the specific places are: 5.1) Spatial Attention Mechanism In the spatial dimension, the influence of different channels on each other is dynamic. By calculating the spatial attention matrix, the attention mechanism is used to dynamically capture this correlation, as shown in equations (16) and (17): in, is the input of the rth spatiotemporal block, C (r-1) is the number of channels of input data, b s ,V s ∈R NxN , is a learnable parameter, σ is an activation function; 5.2) Temporal Attention Mechanism In the time dimension, the data on different time slices of the same channel also have dynamic correlations. The attention mechanism is used to dynamically capture this correlation, as shown in formulas (18) and (19): in, U1∈R N , is a learnable parameter, σ is an activation function; 5.3) Spatial Dependency Modeling Microseismic data and charge data are essentially graph structures. The features of each node are regarded as signals on the graph. Graph convolution based on graph spectrum theory is used to process each time slice. The signal correlation of mine seismic data in the spatial dimension is used to convert the graph into an algebraic form and analyze the topological properties of the graph. The convolution operation of the graph signal is equal to the product of these signals transformed into the spectral domain by the graph Fourier transform. When the scale of the graph is large, the Chebnet model is used, as shown in formula (20). After adding the spatial attention mechanism, it is shown in formula (21): 5.4) Time-dependent modeling After the graph convolution operation captures the adjacent information of each node on the graph in the spatial dimension, it further stacks the standard convolution layer in the time dimension to update the node signal by merging the information on adjacent time slices. After capturing the neighborhood information in the spatial dimension, a standard convolution layer is added in the temporal dimension to update the node information by merging the domain time slice information, as shown in formula (22):

Citation Information

Patent Citations

  • Space-time diagram node attribute prediction method fusing adaptive graph diffusion convolutional network

    CN115828990A

  • Rock burst time sequence prediction model construction method based on small sample learning

    CN115983465A

  • Image processing method based on space-time sequence prediction

    CN118259350A

  • Mine micro-seismic early warning method, device and equipment based on space-time diagram neural network

    CN118915143A

  • Method for optimizing coal mine micro-seismic positioning based on deep learning

    CN118962780A

Cited By

  • Scenic spot flow prediction method, device, equipment, medium and program product

    CN120633956A

  • Mine earthquake intelligent identification method based on multi-mode deep learning and signal processing

    CN120742406A

  • Real-time intelligent prediction method for tunneling thrust of shield tunneling machine based on time sequence deep learning neural network

    CN121525516A

  • Mine micro-seismic multivariable time sequence prediction method based on double-flow structure

    CN121679672A