Mine pressure characteristic parameter extraction method and system based on attention mechanism
Through the method based on attention mechanism, an attention mechanism model is established and the characteristic parameters of ore pressure data are extracted using the self-attention and channel attention modules, the problem of low accuracy of extraction of ore pressure characteristic parameters in the existing technology is solved, and higher extraction accuracy and model expression ability are achieved.
Patent Information
- Application Number
- CN202510024497.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-10
AI Technical Summary
When extracting the ore pressure characteristic parameters of hydraulic support, the prior art is affected by various factors such as the cover rock movement and the working state of the support, resulting in large data fluctuations and low extraction accuracy.
Using an attention mechanism-based method, the ore pressure data set is divided into training samples and test samples by establishing an attention mechanism model, and the characteristic parameters in the ore pressure data are extracted using the self-attention module and the channel attention module.
It improves the accuracy of extraction of mineral pressure feature parameters, can capture high-level abstract features and low-level detailed features in the data more accurately, and enhances the model's expression ability and feature extraction ability.
Smart Images

Figure CN120123720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of coal mines, and in particular to a method and system for extracting characteristic parameters of mine pressure based on an attention mechanism. Background Art
[0002] In an intelligent working face, as a key support device, the working resistance of a hydraulic support reflects the state of the support-surrounding rock, and characteristic parameters such as the initial support force and the final resistance in its data are the core indicators for evaluating the working conditions of the hydraulic support and the roof stability. Accurately extracting characteristic parameters is of crucial significance for underground production safety.
[0003] Currently, the main method for extracting characteristic parameters of mine pressure is the formula extraction method. When the data curve has a regular pattern or stable fluctuations, this method can relatively accurately automatically extract characteristic parameters. However, due to the influence of various factors such as overlying rock movement and the working state of the support, the data of the support pressure often shows large fluctuations, resulting in a relatively low overall extraction accuracy.
[0004] Therefore, there is an urgent need for a method and system that can accurately extract characteristic parameters of mine pressure. Summary of the Invention
[0005] In response to the problems and needs raised above, the present solution proposes a method and system for extracting characteristic parameters of mine pressure based on an attention mechanism. By adopting the following technical features, the above technical objectives can be achieved, and many other technical effects can be brought.
[0006] An object of the present invention is to propose a method for extracting characteristic parameters of mine pressure based on an attention mechanism, including the following steps:
[0007] S10: Collect mine pressure data, and annotate the mine pressure data to obtain a mine pressure data set;
[0008] S20: Divide the mine pressure data set into training samples and test samples;
[0009] S30: Establish an attention mechanism model, input the training samples into the attention mechanism model, extract the characteristic parameters in the mine pressure data set, and verify the attention mechanism model;
[0010] S40: Input the test samples into the trained attention mechanism model to obtain the characteristic parameters of the mine pressure data set.
[0011] In addition, the method for extracting characteristic parameters of mine pressure based on an attention mechanism according to the present invention may further have the following technical features:
[0012] In an example of the present invention, in the step S20, dividing the mine pressure data set into training samples and test samples includes the following steps:
[0013] S21: Slice the strata pressure data according to the frequency of the collected strata pressure data;
[0014] S22: Determine the length and step size of the collected sliding window;
[0015] S23: Save the data under each sliding window, mark the characteristic parameters of the initial support force, final resistance and corrected value of the initial support force therein, and save the marking result at the same time;
[0016] S24: Construct a strata pressure data set and determine the data and its labels.
[0017] In an example of the present invention, in the step S30, the attention mechanism model includes:
[0018] A data preprocessing module, a data encoding module, a self-attention module, a channel attention module and an aggregation output module connected in series in sequence;
[0019] Wherein, the data preprocessing module is configured to perform dimensionality increase processing on the strata pressure data; the data encoding module is configured to extract local features of the time series data and perform normalization processing on the local features in the channel dimension; the self-attention module is configured so that the model can simultaneously capture high-level abstract features and low-level detailed features; the channel attention module is configured to perform adaptive weighting on different feature dimensions to capture and strengthen features of different scales; the aggregation output module is configured to obtain the category information of each data point by fusing the results of the self-attention module and the channel attention module.
[0020] In an example of the present invention, the data preprocessing module includes: a first linear layer configured to perform dimensionality increase processing on the data.
[0021] In an example of the present invention, the data encoding module includes:
[0022] A first convolutional layer, a second linear layer and a first splicing layer, the first convolutional layer is configured to extract local features of the time series data; the second linear layer is configured to map the feature dimension of the input data and adjust the feature dimension without changing the time series length; the first splicing layer is configured to splice the features processed by the first convolutional layer and the second linear layer.
[0023] In an example of the present invention, the data encoding module further includes:
[0024] Two activation function layers, a BatchNorm layer, and a LayerNorm layer. One of the serially-connected activation function layers and the BatchNorm layer are connected to the first convolutional layer, and the other serially-connected activation function layer and the LayerNorm layer are connected to the second linear layer; wherein, the activation function layer is used to introduce non-linearity to the features, the BatchNorm layer is used to normalize the distribution of the features in the batch, and the LayerNorm layer normalizes the features in the channel dimension.
[0025] In an example of the present invention, the self-attention module includes: a first upper part, a first lower part, a second LayerNorm layer, and an FFN layer; wherein, the upper part and the lower part are connected in parallel and serially-connected to the second LayerNorm layer and the FFN layer in sequence. The result obtained by multiplying the output result of the upper part by the feature fusion matrix of the output of the lower part is added to the original input data and then input into the second LayerNorm layer;
[0026] The first upper part includes: a second splicing layer, a second convolutional layer, a second activation function layer, and a third convolutional layer connected in sequence; wherein,
[0027] The second splicing layer is configured to splice multiple replicated data;
[0028] The second convolutional layer is configured to perform convolutional processing on the spliced data;
[0029] The second activation function layer is configured to activate the convolutional output;
[0030] The third convolutional layer is configured to further process the activated data;
[0031] The first lower part includes: three fourth convolutional layers connected in parallel and in sequence, a multi-scale attention dot product module, a fifth convolutional layer, a third activation function layer, and a sixth convolutional layer;
[0032] The three fourth convolutional layers are respectively configured to calculate the query Q, the key K, and the value V; wherein, a position encoding layer is introduced into the fourth convolutional layer for calculating the query Q and the key K to obtain the query Qpos and the key Kpos, and the position encoding layer is configured to capture the position information in the sequence;
[0033] The multi-scale attention dot product module is configured to multiply Qpos, the key Kpos, and the value V to obtain the attention output Attention(Qpos, Kpos, V);
[0034] The fifth convolutional layer is configured to perform convolutional processing on the attention output;
[0035] The third activation function layer is configured to activate the convolutional output;
[0036] The sixth convolutional layer is configured to perform further convolutional processing on the activated data;
[0037] The second LayerNorm layer is configured to normalize the data;
[0038] The FFN layer is configured to perform further non-linear transformation on the normalized data to further enhance the model's expressive ability and feature extraction ability.
[0039] In an example of the present invention, the position encoding layer includes:
[0040] A training parameter module, configured to obtain the position query Q and the position key K by adding the training parameter pos_embed to the query Q and the key K pos and the position key K pos , for capturing the position information in the sequence;
[0041] A score matrix module, configured to multiply the position query Q pos and the position key K pos to obtain an attention score matrix;
[0042] A Softmax function module, configured to convert the score matrix into an attention weight matrix;
[0043] An attention calculation module, configured to multiply the attention weight matrix by the value V to obtain the attention output Attention(Q pos ,K pos ,V), and add the attention output result to V to obtain V res , with the output shape of (B,C,LEN).
[0044] In an example of the present invention, the attention mechanism model further includes:
[0045] A second splicing layer, configured to splice the outputs of the self-attention module and the channel attention module in the channel dimension;
[0046] A third linear layer, configured to reduce the data dimension from 2C to C through linear mapping;
[0047] A residual connection layer, configured to add the spliced output and the input features, and perform downsampling on the residual part;
[0048] A second ReLU activation function layer, for introducing non-linearity to the features;
[0049] A second LayerNorm layer, configured to normalize the features in the channel dimension.
[0050] In one example of the present invention, the channel attention module includes: a second upper part, a second lower part, an activation function layer, and a Dropout layer. Among them, the result obtained by multiplying the second upper part and the second lower part is added to the original data and then input into the activation function layer;
[0051] The second upper part includes: a GroupNorm layer, a first average pooling layer, a seventh convolutional layer, and a second Softmax function layer connected in series in sequence
[0052] The GroupNorm layer is configured to perform grouped normalization processing on the input channels;
[0053] The first average pooling layer is configured to perform pooling in the channel dimension and calculate the average value of each channel;
[0054] The seventh convolutional layer is configured to perform feature extraction and enhancement in a local area;
[0055] The second Softmax function layer is configured to convert the average value of each channel into a weight;
[0056] The eighth convolutional layer is configured to perform a convolutional operation on the channel features processed by the second Softmax function layer;
[0057] The second average pooling layer is configured to perform pooling processing on the output of the eighth convolutional layer;
[0058] The second lower part includes: a first Softmax function layer and a Sigmoid layer connected in series in sequence;
[0059] The first Softmax function layer is configured to perform Softmax processing on the output of the second average pooling layer;
[0060] The Sigmoid layer is configured to map the features to the range of [0, 1];
[0061] The activation function layer is configured to activate the Sigmoid output;
[0062] The Dropout layer is configured to randomly discard the activated features to prevent overfitting;
[0063] The attention calculation layer is configured to multiply the weight by the original channel features and output in the shape of (B, C, LEN).
[0064] Another object of the present invention is to propose an extraction system for mine pressure characteristic parameters based on an attention mechanism, including:
[0065] A data acquisition device, configured to acquire mine pressure data and label the mine pressure data to obtain a mine pressure data set;
[0066] A data division device, configured to divide the mine pressure data set into training samples and test samples;
[0067] A model establishment device, configured to establish an attention mechanism model, input the training samples into the attention mechanism model, extract feature parameters in the mine pressure data set, and verify the attention mechanism model;
[0068] A parameter acquisition device, configured to input the test samples into the trained attention mechanism model to obtain the feature parameters of the mine pressure data set.
[0069] In the following, the optimal embodiments of implementing the present invention will be described in more detail with reference to the accompanying drawings, so that the features and advantages of the present invention can be easily understood. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly introduced below. Among them, the accompanying drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.
[0071] Figure 1 It is a schematic structural diagram of an attention mechanism model according to an embodiment of the present invention;
[0072] Figure 2 It is a schematic structural diagram of a data preprocessing module and a data encoding module according to an embodiment of the present invention;
[0073] Figure 3 It is a schematic structural diagram of a channel attention module according to an embodiment of the present invention;
[0074] Figure 4 It is a schematic structural diagram of a self-attention module according to an embodiment of the present invention;
[0075] Figure 5 It is a schematic diagram of inputting newly acquired mine pressure data into the model to extract feature parameters according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] To make the objectives, technical solutions, and advantages of the technical solutions of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of specific embodiments of the present invention. The same reference numerals in the drawings represent the same components. It should be noted that the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0077] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The "first", "second", and similar terms used in the specification and claims of this patent application for the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "one" do not necessarily denote a quantity limitation. The terms "comprising" or "including" and similar terms mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0078] A method for extracting mine pressure characteristic parameters based on an attention mechanism according to a first aspect of the present invention, as Figure 1 shown, includes the following steps:
[0079] S10: Collect mine pressure data, and label the mine pressure data to obtain a mine pressure data set;
[0080] S20: Divide the mine pressure data set into training samples and test samples;
[0081] S30: Establish an attention mechanism model, input the training samples into the attention mechanism model, extract the characteristic parameters in the mine pressure data set, and verify the attention mechanism model;
[0082] S40: Input the test samples into the trained attention mechanism model to obtain the characteristic parameters of the mine pressure data set.
[0083] In an example of the present invention, in the step S20, dividing the mine pressure data set into training samples and test samples includes the following steps:
[0084] S21: Slice the mine pressure data according to the frequency of the collected mine pressure data;
[0085] S22: Determine the length and step size of the acquisition sliding window;
[0086] S23: Save the data under each sliding window, and label the characteristic parameters of the initial support force, final resistance, and initial support force correction value therein, and at the same time save the labeling result;
[0087] S24: Construct a strata pressure dataset and determine the data and its labels.
[0088] In an example of the present invention, as Figures 2 - 4 shown, in the step S30, the attention mechanism model includes:
[0089] A data preprocessing module, a data encoding module, a self-attention module, a channel attention module, and an aggregation output module connected in series in sequence;
[0090] Among them, the data preprocessing module is configured to perform dimensionality increase processing on the strata pressure data; the data encoding module is configured to extract local features of the time series data and perform normalization processing on the local features in the channel dimension; the self-attention module is configured to enable the model to simultaneously capture high-level abstract features and low-level detailed features; the channel attention module is configured to perform adaptive weighting on different feature dimensions to capture and strengthen features of different scales; the aggregation output module is configured to obtain the category information of each data point by fusing the results of the self-attention module and the channel attention module.
[0091] In an example of the present invention, the data preprocessing module includes: a first linear layer configured to perform dimensionality increase processing on the data.
[0092] For example, the network input is (B*Len), where B is the batch size and Len is the number of strata pressure data in each batch, which is the window size; the data is input into the data preprocessing module, and the data preprocessing module consists of a first linear layer to perform dimensionality increase on the data. After passing through this layer, the output shape is (B*Len*32).
[0093] In an example of the present invention, the data encoding module includes:
[0094] A first convolutional layer, a second linear layer, and a first splicing layer. The first convolutional layer is configured to extract local features of the time series data; the second linear layer is configured to map the feature dimension of the input data and adjust the feature dimension without changing the time series length; the first splicing layer is configured to splice the features processed by the first convolutional layer and the second linear layer.
[0095] In an example of the present invention, the data encoding module further includes:
[0096] Two first ReLU activation function layers, a BatchNorm layer, and a LayerNorm layer. One of the concatenated first ReLU activation function layers and the BatchNorm layer is connected to the first convolutional layer, and the other of the concatenated first ReLU activation function layers and the LayerNorm layer is connected to the second linear layer; wherein, the first ReLU activation function layer is used to introduce non-linearity to the features, the BatchNorm layer is used to normalize the distribution of the features in the batch, and the LayerNorm layer normalizes the features in the channel dimension.
[0097] Specifically, in this data encoding module, the preprocessed data (with shape B*len*32) is first input into the first convolutional layer (Conv1d) and the second linear layer (Linear). The first convolutional layer is mainly used to extract local features of the time series data. By adjusting the number of channels from 32 to 256, it can effectively learn the local patterns of the input data. Its output has a shape of B*256*len.
[0098] The second linear layer directly maps the feature dimension of the input data, adjusting from 32 dimensions to 256 dimensions without changing the time series length. The role of this layer is the transformation of global features without paying attention to the locality of the data in the time dimension. Its output shape is (B*len*256).
[0099] Subsequently, the first ReLU activation function layer and the BatchNorm layer (batch normalization) are applied to the output of the first convolutional layer. The ReLU activation function layer introduces non-linearity to the convolutional features, and the BatchNorm layer normalizes the distribution of the features in the batch, thereby accelerating training and stabilizing the network output. Its output shape remains (B*256*len).
[0100] For the output of the second linear layer, the first ReLU activation function layer is also applied, but the LayerNorm layer (layer normalization) is used. The ReLU activation function layer introduces non-linearity to the features, and the LayerNorm layer normalizes the features in the channel dimension, which helps to eliminate the internal covariate shift problem caused by the increase in the number of layers. Its result shape is (B*len*256). After that, the output of the LayerNorm layer will be transformed into the shape of (B*256*len) for subsequent concatenation operations.
[0101] Finally, the two feature maps processed by the BatchNorm layer and the LayerNorm layer are concatenated (Concat) along the channel dimension (the second dimension) through the first concatenation layer to obtain the final output shape (B * 512 * len). Through this parallel structure, the local features extracted by the first convolutional layer and the global features extracted by the linear layer are effectively fused, providing a richer feature representation and providing multi-scale information input for the subsequent network layers.
[0102] The specific formula is as follows:
[0103] H Conv = σ(W Conv * H g + b Conv )
[0104] H L = W L · H g + b L
[0105] H 1 = Concat(H Conv + H L )
[0106] Among them, H g is the feature after high-dimensional mapping processing, W L and W Conv are the weight matrices of the fully connected layer and the convolutional layer respectively, b L and b Conv are the bias terms, * represents the convolution operation, and σ is the activation function.
[0107] In an example of the present invention, the self-attention module includes: a first upper part, a first lower part, a second LayerNorm layer, and an FFN layer; wherein, the upper part and the lower part are connected in parallel and sequentially connected to the second LayerNorm layer and the FFN layer, and the result obtained by multiplying the output result of the upper part by the feature fusion matrix of the output of the lower part is added to the original input data and then input into the second LayerNorm layer;
[0108] The first upper part includes: a second concatenation layer, a second convolutional layer, a second activation function layer, and a third convolutional layer connected in sequence; wherein,
[0109] The second concatenation layer is configured to concatenate multiple copied data;
[0110] The second convolutional layer is configured to perform convolution processing on the concatenated data;
[0111] The second activation function layer is configured to activate the convolution output;
[0112] The third convolutional layer is configured to further process the activated data;
[0113] The first lower part includes: three fourth convolutional layers connected in parallel and cascaded in sequence, a multi-scale attention dot product module, a fifth convolutional layer, a third activation function layer, and a sixth convolutional layer;
[0114] The three fourth convolutional layers are respectively configured to calculate query Q, key K, and value V; among them, a position encoding layer is introduced in the fourth convolutional layer for calculating query Q and key K to obtain query Qpos and key Kpos, and the position encoding layer is configured to capture the position information in the sequence;
[0115] The multi-scale attention dot product module is configured to multiply Qpos, key Kpos, and value V to obtain an attention output Attention(Qpos, Kpos, V);
[0116] The fifth convolutional layer is configured to perform convolutional processing on the attention output;
[0117] The third activation function layer is configured to activate the convolutional output;
[0118] The sixth convolutional layer is configured to perform further convolutional processing on the activated data;
[0119] The second LayerNorm layer is configured to normalize the data;
[0120] The FFN layer is configured to perform further non-linear transformation on the normalized data to further enhance the model's expression ability and feature extraction ability.
[0121] Specifically, the first upper part includes: a second splicing layer, a second convolutional layer, a second activation function layer, and a third convolutional layer connected in sequence; among them,
[0122] The second splicing layer is configured to splice multiple copied data; the function of this layer is to splice data from different sources (such as query Q, key K, and value V), fuse their information, and provide a richer feature representation for subsequent convolutional processing.
[0123] The second convolutional layer is configured to perform convolutional processing on the spliced data; this convolutional layer is used to perform feature transformation and integration on the spliced data. Through convolutional operations, hidden patterns and features in the data can be extracted to make it suitable for subsequent attention mechanisms.
[0124] The second activation function layer, configured to activate the convolutional output; an activation function (such as ReLU) can introduce a non-linear transformation, enabling the model to learn more complex features. Through activation, the expressive power of the network is enhanced, thereby improving the performance of the model.
[0125] The third convolutional layer, configured to further process the activated data; this convolutional layer is used to further compress or expand the activated data and adjust the dimension of the feature map to make it more suitable for attention calculation or subsequent operations.
[0126] The offline part includes:
[0127] Three fourth convolutional layers, respectively configured to calculate the query Q, key K, and value V; the role of these three convolutional layers is to generate the query Q, key K, and value V, corresponding to Q, K, and V in the self-attention mechanism respectively. Through convolution, the original input is mapped to different feature spaces, generating different feature representations as the basis for calculating attention.
[0128] Among them, a position encoding layer is introduced in the fourth convolutional layer for calculating the query Q and key K to obtain the query Qpos and key Kpos, and the position encoding layer is configured to capture the position information in the sequence;
[0129] The role of the position encoding layer is to add position information to Q and K, enabling the model to capture the relative and absolute positions of elements in the sequence. Position encoding is a very crucial part of sequence data and can help the model understand the temporal relationship and context information.
[0130] The multi-scale attention dot product module, configured to multiply Qpos, key Kpos, and value V to obtain the attention output Attention(Qpos, Kpos, V); the multi-scale attention dot product module calculates the attention output by multiplying Qpos, Kpos, and V. The role of this module is to calculate the attention weights based on the relationship between Q, K, and V and obtain the final attention output through weighted summation. This step is the core part of the self-attention mechanism, allowing the model to dynamically adjust the allocation of attention according to the context information.
[0131] The fifth convolutional layer, configured to perform convolutional processing on the attention output; this convolutional layer is used to further process and extract features from the attention output. Through convolutional operations, the model can further extract higher-level features to prepare for subsequent operations.
[0132] The third activation function layer, configured to activate the convolutional output; the role of the activation function layer is to increase the non-linear characteristics, making the features after convolution more abundant, thereby helping the model capture more complex patterns and features.
[0133] The sixth convolutional layer is configured to perform further convolutional processing on the activated data; this convolutional layer is used to further compress or expand the dimension of the feature map, adjust the feature representation to adapt to subsequent layers or outputs. It continues to optimize the features extracted from the attention output.
[0134] The second LayerNorm layer is configured to perform normalization processing on the data; the role of the second LayerNorm layer is to perform normalization processing on the data. By normalizing the input, it helps to accelerate training, avoid the problems of gradient vanishing or explosion, and improve the stability of the model.
[0135] The FFN layer is configured to perform further feature processing on the normalized data; the FFN layer (feed-forward neural network) is used to perform further non-linear transformation on the features in the final stage. The FFN usually contains multiple fully-connected layers and is processed through activation functions to further enhance the model's expressive ability and feature extraction ability.
[0136] Among them, the output result of the upper-line part is dot-multiplied with the feature fusion matrix of the output of the lower-line part to obtain H21, and H21 is added to the original input data as the input data.
[0137] The outputs of the upper-line part and the lower-line part undergo a feature fusion operation, that is, their feature matrices are dot-multiplied to obtain a new feature matrix H21. Through dot-multiplication, the model combines the feature information of the two parts to obtain a fused representation. Then, this fused feature matrix H21 is added to the original input data to form a residual connection, retaining the feature information of the original input and further enhancing the model's learning ability and stability.
[0138] In an example of the present invention, the position encoding layer includes:
[0139] The training parameter module is configured to add the training parameter pos_embed to the query Q and the key K to obtain the position query Q pos and the position key K pos , which is used to capture the position information in the sequence;
[0140] The score matrix module is configured to multiply the position query Q pos and the position key K pos to obtain the attention score matrix;
[0141] The Softmax function module is configured to convert the score matrix into an attention weight matrix;
[0142] The attention calculation module is configured to multiply the attention weight matrix by the value V to obtain the attention output Attention(Q pos ,K pos,V), and add the attention output result to V to obtain V res , and the output shape is (B, C, LEN).
[0143] Input the result of the data encoding module into the self-attention module (DFA module). The input data shape is (B, C, LEN), where B represents the batch size, C is the number of input channels (feature dimension, usually 512), and LEN is the sequence length. Three second convolutional layers, to_q, to_k, and to_v, are used in the module to calculate the query (Q), key (K), and value (V) respectively. The shape of the input data is (B, C, LEN), and the output shape is the same. Calculate Q, K, and V through the second convolutional layer operation. The convolutional kernel size is 3, and padding = 1 ensures that the output and input dimensions are consistent. After calculating Q, K, and V, they are concatenated. Next, calculate Q, K, and V respectively through the convolutional operation of the second convolutional layer to obtain QKV.
[0144] Next, calculate Q, K, and V respectively through the convolutional operation of the second convolutional layer, and then concatenate Q, K, and V in the channel dimension to form QKV. This operation can integrate the features of the three, facilitating the subsequent multi-head self-attention mechanism.
[0145] To capture the position information in the sequence, the module introduces position encoding (implemented through the trainable parameter pos_embed) and adds it to Q and K; here, the shape of pos_embed is 1×d model ×1, which provides an offset of position information for each position in the sequence, thus retaining the position information of the sequence.
[0146] The position encoding is added to Q and K through the trainable parameter pos_embed to obtain Q pos , K pos , used to capture the position information in the sequence. The shape of pos_embed is (1, d_model, 1). Q pos , K pos After multiplication, calculate the attention score matrix attn_, and the matrix shape is (B, n_head, LEN, LEN). The attention scores are scaled to prevent the values from being too large. Then, convert the score matrix into the attention weight matrix attn through the Softmax function module, and then multiply it by V to obtain the final attention output Attention(Q pos , K pos , V). The final attention output is added to V to obtain the result V, and its shape is B×C×LEN; V res Then, obtain V through the second convolutional layer and the activation function out, the processed result is dot-multiplied with the online feature fusion matrix to obtain H 21 .
[0147] The specific formula steps are as follows:
[0148] The double-layer fusion attention module consists of upper and lower branches. Similarly, a 1x1 convolutional layer more suitable for preserving the spatial structure is used to convert the output of the Backbone into three matrices Q, K, and V.
[0149] Q, K, V = Conv1×1(H 1 )
[0150] Online, through 3x3 convolution, auxiliary local details are effectively aggregated from Q, K, and V.
[0151] QKV = Conv3×3(Concat(Q, K, V))
[0152] Subsequently, a linear projection with an activation function and batch normalization is used to compress the dimension (2C_qk + C_v) to C to generate detail enhancement weights, which can deeply understand the initial support force and the final resistance.
[0153] W s = BN(ReLU(Conv1×1(QKV)))
[0154] In the offline processing, first, position encoding is performed on the Q and K matrices:
[0155] Q pos , K pos = PosEncoding(Q, K)
[0156] Then, it is combined with the V matrix to perform multi-head self-attention calculation.
[0157]
[0158] Next, the V matrix with spatial attention weights is residually connected to the V before attention calculation, and further processed through a 1x1 convolution and a sigmoid activation function.
[0159] V res = Attention(Q pos , K pos , V) + V
[0160] V out = σ(Conv1×1(V res ))
[0161] The processed result is dot - multiplied with the feature fusion matrix on the production line. Finally, these weight matrices adjusted by spatial attention are multiplied by the features of the QKV fusion matrix processed by 3x3 convolution.
[0162] H 21 = V out ⊙W s
[0163] In an example of the present invention, the attention mechanism model further includes:
[0164] A second splicing layer, configured to splice the outputs of the self - attention module and the channel - attention module in the channel dimension;
[0165] A third linear layer, configured to reduce the data dimension from 2C to C through linear mapping;
[0166] A residual connection layer, configured to add the spliced output and the input features, and downsample the residual part;
[0167] A second ReLU activation function layer, used to introduce non - linearity to the features;
[0168] A second LayerNorm layer, configured to normalize the features in the channel dimension.
[0169] Specifically, when combining the DFA and MSFA modules, first, the outputs of the self - attention module (DFA module) and the channel - attention module (MSFA module) are respectively obtained. The output shapes of both are B*C*LEN, where B is the batch size, C is the number of channels, and LEN is the sequence length.
[0170] The second splicing layer splices the outputs of the two modules in the channel dimension to form a comprehensive representation containing self - attention and channel - attention features. The shape after splicing is B*2C*LEN. This splicing operation fuses the self - attention and channel - attention information together, enriching the feature representation.
[0171] The spliced data passes through the third linear layer linear1 to reduce the dimension from 2C to C, so as to reduce the computational complexity and focus on important features. This step, through linear mapping, not only maintains the representation ability of the original features but also eliminates the redundant information brought by the extra channels, making the subsequent calculations more efficient.
[0172] To further enhance the stability of the model, a residual connection layer is introduced to the spliced output to add the spliced output and the input features. The residual part is downsampled by the downsample module to ensure that the dimension matches B*C*LEN. The downsample module reduces the redundant channels through appropriate dimensionality reduction operations (such as convolution or linear projection) to ensure consistent dimensions when adding the residuals.
[0173] Finally, the second ReLU activation function layer is introduced to introduce non-linearity and deepen the expressive ability of feature representation. Then, the second LayerNorm layer normalizes the features in the channel dimension to further stabilize the model output and reduce the problem of internal covariate shift. The final output shape is B*C*LEN, which contains the information combining self-attention and channel attention and has undergone reasonable dimensionality reduction and normalization to ensure the diversity and stability of feature information.
[0174] In an example of the present invention, the channel attention module includes: a second upper part, a second lower part, an activation function layer, and a Dropout layer, wherein the result obtained by multiplying the second upper part and the second lower part is added to the original data and then input into the activation function layer;
[0175] The second upper part includes: a GroupNorm layer, a first average pooling layer, a seventh convolutional layer, and a second Softmax function layer connected in series in sequence;
[0176] The GroupNorm layer is configured to perform grouped normalization processing on the input channels;
[0177] The first average pooling layer is configured to perform pooling in the channel dimension and calculate the average value of each channel;
[0178] The seventh convolutional layer is configured to perform feature extraction and enhancement in a local area;
[0179] The second Softmax function layer is configured to convert the average value of each channel into weights;
[0180] The eighth convolutional layer is configured to perform a convolutional operation on the channel features processed by the second Softmax function layer;
[0181] The second average pooling layer is configured to perform pooling processing on the output of the eighth convolutional layer;
[0182] The second lower part includes: a first Softmax function layer and a Sigmoid layer connected in series in sequence;
[0183] The first Softmax function layer is configured to perform Softmax processing on the output of the second average pooling layer;
[0184] The Sigmoid layer is configured to map the features to the range of [0,1];
[0185] The activation function layer is configured to activate the Sigmoid output;
[0186] The Dropout layer is configured to randomly discard the activated features to prevent overfitting;
[0187] The attention calculation layer is configured to multiply the weights with the original channel features and output in the shape of (B, C, LEN).
[0188] Specifically, the second upper part includes: a GroupNorm layer, a first average pooling layer, a seventh convolutional layer, and a second Softmax function layer connected in series in sequence;
[0189] The GroupNorm layer is configured to perform grouped normalization processing on the input channels;
[0190] The function of this layer is to perform grouped normalization processing on the input channel data. By dividing the input channels into several groups and performing normalization within each group, the GroupNorm layer can avoid the influence of the batch size on the normalization process, especially suitable for the training of small batch data. In this way, the model can be trained more stably and accelerate convergence.
[0191] The first average pooling layer is configured to perform pooling in the channel dimension and calculate the average value of each channel;
[0192] This average pooling layer performs pooling on the input data in the channel dimension and calculates the average value of each channel. Through the pooling operation, the model can extract global statistical information from each channel, and these information can help the model capture the global features and context of the data, thereby enhancing the mutual relationship between channels.
[0193] The seventh convolutional layer is configured to perform feature extraction and enhancement in the local area;
[0194] The seventh convolutional layer extracts features in the local area through convolutional operations and enhances the local representation ability of the input data. The convolutional layer can help the model learn spatial local patterns, enabling the features in the channel dimension to be further extracted and enhanced. This provides important local feature information for subsequent channel attention calculation.
[0195] The second Softmax function layer is configured to convert the average value of each channel into weights;
[0196] The second Softmax function layer normalizes the average value of each channel, making the weight value of each channel between [0, 1], and the sum of the weights of all channels is 1. Through Softmax, the model can distribute attention among all channels, thereby controlling which channels contribute more to the output. This can dynamically adjust the weights of each channel and improve the selective expression ability of the model.
[0197] The eighth convolutional layer is configured to perform a convolution operation on the channel features processed by Softmax;
[0198] The eighth convolutional layer performs a convolution operation on the weighted features obtained after Softmax processing. The convolution operation helps to extract higher-order patterns from the weighted channel features and further optimize the feature representation of each channel. Through convolution, the model can learn richer feature representations and pass them to the next layer for further processing.
[0199] The second average pooling layer is configured to perform pooling on the output of the eighth convolutional layer;
[0200] The second average pooling layer performs pooling on the output of the eighth convolutional layer in the channel dimension to further extract the global information of each channel and help the model better understand the importance of different channels. The pooling operation can reduce the feature dimension while retaining key information, thereby improving the computational efficiency and avoiding overfitting.
[0201] The first Softmax function layer is configured to perform Softmax processing on the output of the second average pooling layer;
[0202] The first Softmax function layer performs Softmax processing on the output of the second average pooling layer, normalizing the pooled channel weights. In this way, the model can dynamically allocate weights between channels, enabling the output of each channel to be weighted according to its importance. The Softmax function ensures a smooth distribution of channel weights, thereby enhancing the model's representation ability.
[0203] The Sigmoid layer is configured to map the features to the range [0, 1];
[0204] The role of the Sigmoid layer is to compress the features so that their value range is between [0, 1]. This step helps the model control the activation degree of each channel by mapping the output of each channel to a probability value, further strengthening the selective attention between channels.
[0205] The activation function layer is configured to activate the Sigmoid output;
[0206] The activation function layer (such as ReLU) helps the model learn more complex features by adding non-linearity. Through activation, the model can capture non-linear patterns in the data and improve the model's fitting ability for complex relationships. The activation layer is usually used to introduce more non-linear expressions, enabling the model to have stronger feature representation capabilities.
[0207] The Dropout layer is configured to randomly discard the activated features to prevent overfitting;
[0208] The Dropout layer randomly drops some neurons during the training process to prevent the model from overfitting to the training data. With Dropout, the model can better generalize, thus improving its performance on unseen data. Dropout helps enhance the robustness of the model and avoid over-reliance on certain specific features during training.
[0209] The attention calculation layer is configured to multiply the weights with the original channel features, and the output shape is (B, C, LEN);
[0210] The attention calculation layer multiplies the weights of each channel with the original channel features to obtain the weighted channel representation. In this way, the model can dynamically adjust its output according to the importance of each channel, making the key channels contribute more to the final result. The output generated by this layer has the shape of (B, C, LEN), maintaining the same dimensional structure as the input data.
[0211] Specifically, in the channel attention module (MSFA module), first, the result of the data encoding module (with the shape of B*C*LEN) is input into the GroupNorm layer to perform group normalization on the input channels. The number of groups n_group is determined by d_model / / group_dim. Through group normalization, the feature distribution within each group is smoothed, which helps to stabilize the training process and avoid gradient instability caused by excessive differences in channel dimension features.
[0212] Next, the average pooling layer (avgpool) pools each feature in the channel dimension to calculate the average value of each channel, thereby compressing the length dimension. The output shape after pooling is B*C*1, which represents the average feature of each channel in the sequence length dimension.
[0213] Then, the pooled channel features are enhanced through a convolution operation. Here, the kernel size of the convolution is 3, and padding = 1 is set to ensure that the output shape is the same as the input. This convolution operation is used for feature extraction and enhancement in the local area, laying the foundation for the subsequent generation of attention weights. The shape of the convolution output remains B*C*1.
[0214] After that, the result of the convolution operation is transformed into weights through the Softmax function layer, ensuring that the weight values are normalized in the channel dimension. Then, the attention calculation layer multiplies the obtained weights with the original channel features (with the shape of B*C*LEN) element by element, so that the feature values of each channel are adjusted according to the weights generated by pooling and convolution. The final output shape still remains B*C*LEN, achieving multi-scale aggregation and dynamic adjustment of features.
[0215] This mechanism allows the MSFA module to perform adaptive weighting on different feature dimensions, enabling it to capture and enhance features of different scales, thereby improving the model's performance in processing sequence data.
[0216] The specific formula steps are as follows:
[0217] In the multi-scale fusion attention module, the output data of the data encoding module is grouped by features to determine the number of features each group is responsible for. H 1 After grouping, the feature H is obtained grou . The grouped feature matrix undergoes 3x3 convolution in the high-scale space for feature extraction to obtain high-dimensional features, resulting in I 1,1 .
[0218] I 1,1 = Conv3×3(H grou )
[0219] Meanwhile, the grouped features are subjected to global average pooling and multiplied by the original features to obtain the channel attention features I in the low-scale space 2,1 , the matrix I 2,1 is then subjected to global average pooling and Softmax to obtain I 1,2 , I 1,1 is also subjected to global average pooling and Softmax to obtain I 2,2 .
[0220] I 2,1 = AvgPool(H grou )·H grou
[0221] I 1,2 = softmax(Avgpool(I 2,1 ))
[0222] I 2,2 = softmax(Avgpool(I 1,1 ))
[0223] Multiply I 1,1 and I 1,2 using Matmal dot product to obtain the feature information Z that combines channel attention and convolution in the high-scale space Ⅰ .
[0224] I 2,1 and I 2,2 are multiplied using Matmul to obtain the feature information Z that combines the advanced feature weights extracted by convolution and channel attention in the low-scale space Ⅱ .
[0225]
[0226]
[0227] Add the high-scale spatial information and the low-scale spatial feature information.
[0228] H 22 =Z Ⅰ +Z Ⅱ
[0229] An extraction system for mining pressure characteristic parameters based on an attention mechanism according to a second aspect of the present invention includes:
[0230] A data acquisition device configured to collect mining pressure data and label the mining pressure data to obtain a mining pressure data set;
[0231] A data division device configured to divide the mining pressure data set into training samples and test samples;
[0232] A model establishment device configured to establish an attention mechanism model, input the training samples into the attention mechanism model, extract the characteristic parameters in the mining pressure data set and verify the attention mechanism model;
[0233] A parameter acquisition device configured to input the test samples into the trained attention mechanism model to obtain the characteristic parameters of the mining pressure data set.
[0234] Specific case
[0235] To verify the effectiveness of the method of the present invention, characteristic parameters are extracted using the coal mine support pressure data in the Shendong mining area. The data set is selected from the underground support resistance data of the coal mines in the Shendong mining area, and specifically, the moment data of the night shift within 3 days is selected. The data collection in the database follows the principle of not storing unchanged data, and a total of about 20,000 pieces of data are collected.
[0236] Since the data has strong front-back correlation, we use a sliding window and a specific step size to construct the data set. According to different data set schemes, the divided data amounts are also different. Two data sets are designed according to the coal mines. The Booltai data set is denoted as DataSet1, and the Yujialiang data set is denoted as DataSet2. Among them, DataSet1 has a total of 21,678 pieces of data, and DataSet2 has a total of 22,452 pieces of data. 200, 400, 600, 800, 1000, 1200 and 10, 15, 20 are respectively used as the window size and the sliding step size for data grouping. After building the data set, we divide the data into a training set, a validation set and a test set according to the ratio of 15:3:2. Among them, the test environment list is shown in Table 1.
[0237] Table 1 is the list of experimental environments
[0238]
[0239] The initial training round is set to 100, the training batch size is set to 32, the learning rate is 0.0001, and NLLLoss is selected as the loss function, and Adam is used as the optimizer for experiments.
[0240] The specific detailed steps are as follows:
[0241] Step 1: Collect electro-hydraulic control support pressure (mine pressure data) data and clean the mine pressure data;
[0242] Step 2: Annotate the mine pressure data and make a data set;
[0243] Step 2.1: Slice the data according to the frequency of the collected mine pressure data;
[0244] Step 2.2: Determine the sliding window length (Window_size) and the step size (Step);
[0245] Step 2.3: Save the data under each sliding window, and annotate the characteristic parameters such as the initial support force, the final resistance, and the corrected value of the initial support force, and save the annotation results at the same time;
[0246] Step 2.4: Build a data set and determine the data and its labels.
[0247] Step 3: Establish an attention model to extract the characteristic parameters in the mine pressure data
[0248] Step 3.1: The network input is (B*Len), where B is the batch size and Len is the number of mine pressure data in each batch, which is the window size.
[0249] Step 3.2: Input the data into the data preprocessing module. The data preprocessing module consists of Linear layers to increase the dimension of the data. After passing through this layer, the output shape is (B*Len*32)
[0250] Step 3.3: Input the results of the data preprocessing module into the data encoding module. The data encoding module consists of 1 Conv1d layer, 1 Linear layer, 2 Relu layers, 1 LayerNorm layer, and 1 BatchNorm layer. The input shape of the data encoding module is B*Len*32, which is input into the Conv1d layer and the Linear layer simultaneously. The output of the Conv1d layer is B*256*Len, and the output of the Linear layer is B*Len*256. Then, input the output of the Conv1d layer into the Relu layer and the BatchNorm layer to obtain the result B*256*Len. Input the output of the Linear layer into the Relu layer and the LayerNorm layer to obtain the result B*Len*256. Finally, transform the output of LayerNorm to B*256*Len, and concatenate the two parts of the results to get B*512*Len. The specific formula is as follows:
[0251] H Conv =σ(W Conv *H g +b Conv )
[0252] H L =W L ·H g +b L
[0253] H 1 =Concat(H Conv +H L )
[0254] Among them, H g is the feature after high-dimensional mapping processing, W L and W Conv are the weight matrices of the fully connected layer and the convolutional layer respectively, b L and b Conv are the bias terms, * represents the convolution operation, and σ is the activation function.
[0255] Step 3.4: Input the result of the data encoding module into the DFA module. The shape of the input data is (B, C, LEN), where B represents the batch size, C is the number of input channels (feature dimension, usually 512), and LEN is the sequence length. Three convolutional layers, to_q, to_k, and to_v, are used in the module to calculate the query (Q), key (K), and value (V) respectively. The shape of the input data is (B, C, LEN), and the output shape is the same. Q, K, and V are calculated through the Conv1d convolution operation. The kernel size of the convolution is 3, and padding = 1 ensures that the output and input dimensions are consistent. After Q, K, and V are calculated, they are concatenated and sent to the subsequent convolution operation to obtain QKV. The positional encoding is added to Q and K through the trainable parameter pos_embed to obtain Q pos , K pos , used to capture the positional information in the sequence. The shape of pos_embed is (1, d_model, 1). Q pos , K pos After multiplication, the attention score matrix attn_ is calculated. The shape of the matrix is (B, n_head, LEN, LEN). The attention scores are scaled to prevent the values from being too large. The scores are transformed into the attention weight matrix attn through the Softmax function and multiplied by V to obtain the final attention output Attention(Q pos , K pos , V), and the result is added to V to obtain V res , and the output shape is (B, C, LEN). V res Then, it passes through the convd1 layer and the activation function to obtain V out , and the processed result is dot-multiplied with the feature fusion matrix of the upper line to obtain H 21 .
[0256] The specific formula steps are as follows:
[0257] The double-layer fusion attention module consists of upper and lower branches. Similarly, a 1x1 convolutional layer more suitable for preserving the spatial structure is used to convert the output of the Backbone into three matrices, Q, K, and V.
[0258] Q, K, V = Conv1×1(H 1 )
[0259] The upper line effectively aggregates auxiliary local details from Q, K, and V through 3x3 convolution.
[0260] QKV = Conv3×3(Concat(Q, K, V))
[0261] Subsequently, a linear projection with an activation function and batch normalization is adopted to compress the dimension (2C_qk + C_v) to C to generate detail enhancement weights, enabling a deep understanding of the initial support force and the final resistance force.
[0262] W s = BN(ReLU(Conv1×1(QKV)))
[0263] In the offline processing, first, positional encoding is performed on the Q and K matrices:
[0264] Q pos , K pos = PosEncoding(Q, K)
[0265] Then, it is combined with the V matrix to perform multi-head self-attention calculation.
[0266]
[0267] Next, the V matrix with spatial attention weights is subjected to residual connection with the V before attention calculation, and further processed through 1x1 convolution and sigmoid activation function.
[0268] V res = Attention(Q pos , K pos , V) + V
[0269] V out = σ(Conv1×1(V res ))
[0270] The processed result is subjected to a dot product operation with the feature fusion matrix in the online process. Finally, these weight matrices adjusted by spatial attention are multiplied by the features of the QKV fusion matrix processed by 3x3 convolution.
[0271] H 21 = V out ⊙W s
[0272] Step 3.5: Input the result of the data encoding module into the MSFA module. The shape of the input data is (B, C, LEN). The GroupNorm layer in the module performs group normalization on the input channels. The specific number of groups is determined by n_group, and n_group is equal to d_model / / group_dim. Then, the avgpool average pooling layer performs pooling in the channel dimension to calculate the average value of each channel, and the output shape is (B, C, 1). The pooled channel features are enhanced through a convolution operation. The kernel size of the convolution is 3, and padding = 1 ensures that the input and output dimensions are consistent. Then, the pooled output undergoes a Softmax operation to be transformed into weights, and the weights are multiplied by the original channel features, and the output shape remains (B, C, LEN).
[0273] The specific formula steps are as follows:
[0274] In the multi-scale fusion attention module, the output data of the data encoding module is grouped by features to determine the number of features responsible for each group. H 1 After grouping, the feature H is obtained grou . The grouped feature matrix undergoes a 3x3 convolution in the high-scale space for feature extraction to obtain high-dimensional features, and I is obtained 1,1 .
[0275] I 1,1 = Conv3×3(H grou )
[0276] At the same time, the grouped features are globally averaged pooled and multiplied by the original features to obtain the channel attention feature I in the bottom-scale space 2,1 , the matrix I 2,1 is then globally averaged pooled and Softmaxed to obtain I 1,2 , I 1,1 is also globally averaged pooled and Softmaxed to obtain I 2,2 .
[0277] I 2,1 = AvgPool(H grou )·H grou
[0278] I 1,2 = softmax(Avgpool(I 2,1 ))
[0279] I 2,2 = softmax(Avgpool(I 1,1 ))
[0280] Multiply I 1,1 with I 1,2Perform a Matmal dot product to obtain the feature information Z that combines channel attention and convolution in the high-scale space Ⅰ 。
[0281] I 2,1 Perform a Matmul multiplication with I 2,2 to obtain the feature information Z that combines the high-level feature weights extracted by convolution and channel attention in the low-scale space Ⅱ 。
[0282]
[0283]
[0284] Add the high-scale space information and the low-scale space feature information together.
[0285] H 22 =Z Ⅰ +Z Ⅱ
[0286] Step 3.6: Combination of the DFA and MSFA modules. The output shapes of both the self-attention module and the channel-attention module are (B, C, LEN). Concatenate the outputs of the two modules to obtain data of (B, 2C, LEN). The concatenated data passes through the linear layer linear1 to reduce the dimension from 2C to C. To ensure the stability of the model, the output after concatenation passes through a residual connection, and the residual part is downsampled by the downsample module to ensure dimension matching. Finally, the output passes through ReLU activation and LayerNorm normalization to obtain the final output with a shape of (B, C, LEN).
[0287] Step 3.7: Input the result after the fusion attention feature extraction module into the aggregation output module to output the class information of each data point, including whether it is the initial support force, the final resistance, the initial support force correction value, etc.
[0288] Step 3.8: Train and validate the model, and save the model weight file with excellent test results.
[0289] Step 4: Input the newly collected mine pressure data into the model to extract feature parameters. The results are as Figure 5 shown:
[0290] Table 2 shows the list of accuracy rates corresponding to different methods
[0291]
[0292] Through the integration of two attention mechanisms, DFA and MSFA, a large number of comparative experiments are shown in Table 2. The model performs outstandingly on different datasets. The accuracy rate of Dataset 1 reaches 98.9%, the precision rate is 99.3%, and the mAP reaches 94.93%. The accuracy rate of Dataset 2 reaches 99.5%, the precision rate is 99.8%, and the mAP is as high as 99.97%. The experiments prove that the model we proposed is more efficient and accurate in the time-series feature classification task. The model comprehensively focuses on the time and space characteristics of the mine pressure data, improving the robustness in different mine environments.
[0293] In the above text, the exemplary embodiments of the method and system for extracting mine pressure characteristic parameters based on the attention mechanism proposed by the present invention are described in detail with reference to the preferred embodiments. However, those skilled in the art can understand that, without departing from the concept of the present invention, various modifications and variations can be made to the above specific embodiments, and various combinations of the technical features and structures proposed by the present invention can be made, without exceeding the protection scope of the present invention. The protection scope of the present invention is determined by the appended claims.
Claims
1. A method for extracting mine pressure characteristic parameters based on attention mechanism, characterized in that: The steps include: S10: Collecting mine pressure data, and labeling the mine pressure data to obtain a mine pressure data set; S20: Divide the mine pressure data set into training samples and test samples; S30: Establishing an attention mechanism model, inputting the training samples into the attention mechanism model, extracting characteristic parameters from the mine pressure data set and verifying the attention mechanism model; S40: Input the test sample value into the trained attention mechanism model to obtain the characteristic parameters of the mine pressure data set.
2. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 1 is characterized in that: In step S20, dividing the mine pressure data set into training samples and test samples includes the following steps: S21: Slicing the mine pressure data according to the frequency of the collected mine pressure data; S22: Determine the length and step size of the acquisition sliding window; S23: saving the data in each sliding window, and marking the characteristic parameters of the initial support force, final resistance and initial support force correction value, and saving the marking results; S24: Construct a mine pressure dataset and determine the data and its labels.
3. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 1 is characterized in that: In step S30, the attention mechanism model includes: A data preprocessing module, a data encoding module, a self-attention module, a channel attention module and an aggregation output module are connected in series in sequence; Among them, the data preprocessing module is configured to perform dimensionality increase processing on the mine pressure data; the data encoding module is configured to extract local features of the time series data and normalize the local features in the channel dimension; the self-attention module is configured so that the model can simultaneously capture high-level abstract features and low-level detail features; the channel attention module is configured to perform adaptive weighting on different feature dimensions to capture and enhance features of different scales; the aggregation output module is configured to obtain the category information of each data point by fusing the results of the self-attention module and the channel attention module.
4. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 3 is characterized in that: The data encoding module comprises: A first convolutional layer, a second linear layer and a first splicing layer, wherein the first convolutional layer is configured to extract local features of time series data; the second linear layer is configured to map the feature dimensions of the input data and adjust the feature dimensions without changing the time series length; the first splicing layer is configured to splice the features processed by the first convolutional layer and the second linear layer.
5. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 4 is characterized in that: The data encoding module also includes: Two first activation function layers, a BatchNorm layer and a LayerNorm layer, one of the first activation function layers and the BatchNorm layer connected in series is connected to the first convolutional layer, and the other first activation function layer and the LayerNorm layer connected in series is connected to the second linear layer; wherein the first activation function layer is used to introduce nonlinearity to the features, the BatchNorm layer is used to normalize the distribution of the features in the batch, and the LayerNorm layer normalizes the features in the channel dimension.
6. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 1 is characterized in that: The self-attention module includes: a first upper line part, a first lower line part, a second LayerNorm layer and an FFN layer; wherein the upper line part and the lower line part are connected in parallel and the second LayerNorm layer and the FFN layer are connected in series in sequence, and the output result of the upper line part is dot-multiplied with the feature fusion matrix output by the lower line part, and the result obtained by adding the original input data is input into the second LayerNorm layer; The first online part includes: a second concatenation layer, a second convolution layer, a second activation function layer, and a third convolution layer connected in series; wherein, The second splicing layer is configured to splice the copied multiple data; The second convolutional layer is configured to perform convolution processing on the concatenated data; The second activation function layer is configured to activate the convolution output; The third convolutional layer is configured to further process the activated data; The first offline part includes: three fourth convolutional layers, a multi-scale attention point multiplication module, a fifth convolutional layer, a third activation function layer and a sixth convolutional layer connected in series and in parallel; The three fourth convolutional layers are respectively configured to calculate the query Q, the key K and the value V; wherein a position encoding layer is introduced in the fourth convolutional layer for calculating the query Q and the key K to obtain the query Qpos and the key Kpos, and the position encoding layer is configured to capture the position information in the sequence; The multi-scale attention point multiplication module is configured to multiply Qpos, key Kpos and value V to obtain the attention output Attention(Qpos, Kpos, V); The fifth convolutional layer is configured to perform convolution on the attention output; The third activation function layer is configured to activate the convolution output; The sixth convolutional layer is configured to perform further convolution processing on the activated data; The second LayerNorm layer is configured to normalize the data; The FFN layer is configured to perform further nonlinear transformation on the normalized data to further enhance the expression and feature extraction capabilities of the model.
7. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 6 is characterized in that: The position encoding layer comprises: The training parameter module is configured to obtain the position query Q by adding the training parameter pos_embed to the query Q and key K pos and position key K pos , used to capture position information in the sequence; The score matrix module is configured to convert the position query Q pos and position key K pos Perform multiplication to obtain the attention score matrix; Softmax function module, configured to convert the score matrix into an attention weight matrix; The attention calculation module is configured to multiply the attention weight matrix by the value V to obtain the attention output Attention(Q pos ,K pos ,V), and add the attention output result to V to get V res , the output shape is (B,C,LEN).
8. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 6 is characterized in that: The attention mechanism model also includes: A second concatenation layer is configured to concatenate the outputs of the self-attention module and the channel attention module in the channel dimension; The third linear layer is configured to reduce the data dimension from 2C to C through linear mapping; The residual connection layer is configured to add the concatenated output and input features and downsample the residual part; The second ReLU activation function layer is used to introduce nonlinearity to the features; The second LayerNorm layer is configured to normalize the features in the channel dimension.
9. The method for extracting mine pressure characteristic parameters based on the attention mechanism according to claim 1 is characterized in that: The channel attention module includes: a second upper line part, a second lower line part, an activation function layer and a Dropout layer, wherein the result obtained by dot multiplication of the second upper line part and the second lower line part is added to the original data and then input into the activation function layer; The second online part includes: a GroupNorm layer, a first average pooling layer, a seventh convolutional layer, and a second Softmax function layer connected in series. The GroupNorm layer is configured to perform group normalization on the input channels; The first average pooling layer is configured to perform pooling in the channel dimension and calculate the average value of each channel; The seventh convolutional layer is configured to perform feature extraction and enhancement in a local area; The second Softmax function layer is configured to convert the average value of each channel into a weight; The eighth convolutional layer is configured to perform a convolution operation on the channel features processed by the second Softmax function layer; a second average pooling layer, configured to perform pooling on the output of the eighth convolutional layer; The second offline part includes: a first Softmax function layer and a Sigmoid layer connected in series; The first Softmax function layer is configured to perform Softmax processing on the output of the second average pooling layer; Sigmoid layer, configured to map features to the range of [0,1]; The activation function layer is configured to activate the Sigmoid output; Dropout layer, configured to randomly discard activated features to prevent overfitting; The attention calculation layer is configured to multiply the weights with the original channel features, and the output shape is (B, C, LEN).
10. A system for extracting mine pressure characteristic parameters based on attention mechanism, characterized in that: include: A data acquisition device is configured to collect mine pressure data and annotate the mine pressure data to obtain a mine pressure data set; A data partitioning device configured to partition a mine pressure data set into training samples and test samples; A model building device, configured to build an attention mechanism model, input the training sample into the attention mechanism model, extract characteristic parameters from the mine pressure data set and verify the attention mechanism model; The parameter acquisition device is configured to input the test sample value into the trained attention mechanism model to obtain the characteristic parameters of the mine pressure data set.