Method for analyzing borehole strain data based on improved Pyraformer neural network

Through the improved Pyraformer neural network and Delta method, the problem of difficult to capture the long-term dependence and insufficient parallelization capabilities of earthquake data in the prior art is solved, and faster model training and higher result reliability are achieved.

CN119961729APending Publication Date: 2025-05-09HAINAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056226.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the long-term dependence of earthquake data when processing drilling strain data. The local receptive field of the convolution kernel limits the network's acquisition of global information, and the parallelization ability is not strong when processing massive data, resulting in a long training time for network models.

Method used

The drilling strain data was analyzed using an improved Pyraformer neural network, which trained and predicted through stacked encoding and decoding layers, and constructed confidence intervals for prediction results in combination with the Delta method to determine the anomaly data.

Benefits of technology

The improved Pyraformer network improves model training speed, can effectively capture global information and avoid gradient vanishing problems, and use the Delta method to improve the reliability of the results and reduce potential uncertainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961729A_ABST
    Figure CN119961729A_ABST
Patent Text Reader

Abstract

The invention discloses a borehole strain data analysis method based on an improved Pyraformer neural network, and the method comprises the steps: obtaining to-be-analyzed station surface strain data, and carrying out the preprocessing of the to-be-analyzed station surface strain data; inputting the preprocessed station surface strain data into an improved Pyraformer neural network for processing, and outputting a prediction analysis result; and constructing a confidence interval of a prediction result for the prediction analysis result by adopting a Delta method, and determining abnormal data according to the confidence interval. According to the method, the model training speed is increased by using the Pyraformer network, and the reliability of the result is improved by using the Delta method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of data analysis, and in particular to an analysis method for drilling strain data based on an improved Pyraformer neural network. Background Art

[0002] Earthquakes, as one of the most destructive natural disasters, have brought immeasurable harm to human society, often causing a large number of casualties and property losses. Compared with other disasters, earthquakes have no special mechanism of occurrence and are more destructive.

[0003] Earthquakes occur due to the movement of crustal plates and the evolution of the earth's internal structure, and their breeding and triggering mechanisms are extremely complex. Although scientists have invested a lot of research resources in this field, they still face many technical challenges and limitations in actual earthquake prediction experiments. In recent years, countries have successively established earthquake early warning systems. The main method of earthquake early warning is based on the monitoring and analysis of earthquake precursor signals. In earthquake prediction research, scientists use various technical means and tools to collect, analyze and interpret earthquake precursor signals. Borehole strain gauges have become one of the important tools for crustal deformation observation due to their high precision, wide bandwidth and good stability.

[0004] With the further improvement and optimization of borehole strain technology, the deployment rate of borehole strain gauges in my country has been greatly improved, and the accumulation of data has also increased. Traditional signal processing methods have shown many shortcomings in the processing of massive data. The use of deep learning methods can effectively solve this problem. Traditional deep learning models have achieved certain results in seismic data analysis, but these methods are difficult to effectively capture the long-term dependencies of seismic data, and the local receptive field of the convolution kernel limits the network's acquisition of global information. In addition, some deep learning methods have weak parallelization capabilities when processing massive data, resulting in a long network model training time. The Pyraformer neural network can effectively capture global information, avoid the gradient vanishing problem, and has strong parallelization capabilities, which greatly improves the speed of network model training. The Delta method is used to construct confidence intervals to improve the reliability of the results and reduce potential uncertainties.

[0005] So far, there has been no report on the use of Pyraformer neural network method and Delta method to analyze drilling strain data. Summary of the invention

[0006] In order to solve the technical problems existing in the above-mentioned prior art, the present invention proposes an analysis method for drilling strain data based on an improved Pyraformer neural network.

[0007] To achieve the above object, the present invention provides a method for analyzing drilling strain data based on an improved Pyraformer neural network, comprising:

[0008] Acquiring station surface strain data to be analyzed, and preprocessing the station surface strain data to be analyzed;

[0009] The pre-processed station surface strain data is input into the improved Pyraformer neural network for processing, and the prediction analysis results are output; wherein the improved Pyraformer neural network is obtained by training and predicting the training set through stacked encoding layers and decoding layers, and the training set is the station surface strain data set;

[0010] The Delta method is used to construct a confidence interval of the prediction result for the prediction analysis result, and abnormal data is determined based on the confidence interval.

[0011] Preferably, the station surface strain data to be analyzed is preprocessed, including:

[0012] Adaptive variational mode decomposition method AVMD is used to decompose the station surface strain data to be analyzed and calculate the initial window mean;

[0013] Slide the station surface strain data to be analyzed in windows, and judge the mean of each window by the initial window mean. If there is abnormal day data, adjust the window length to decompose the abnormal day data separately, and perform VMD decomposition on the adjusted window;

[0014] The trend terms and intrinsic mode function components obtained by the adaptive variational mode decomposition method AVMD are sorted, the solid tide is decomposed, the annual trend and solid tide components of the data are removed, and the preprocessed station surface strain data are obtained.

[0015] Preferably, the initial window mean is calculated as follows:

[0016]

[0017] In the formula, m prev represents the initial window mean; L represents the initial window length; x[k] represents the current data sequence number;

[0018] The mean value of each window is judged by the initial window mean value, specifically:

[0019] m i >2*m prev ;

[0020]

[0021] In the formula, mi Indicates the mean of the window; m prev represents the initial window mean; L represents the initial window length; k, n represent the current data sequence number respectively; m i,d represents the mean of one day in the window; D represents the data points of each day; d represents the day number in the window.

[0022] Preferably, the preprocessed station surface strain data is input into the improved Pyraformer neural network for processing, including:

[0023] The data after AVMD decomposition is aggregated at different scales through the coarse-scale construction module CSCM, a multi-resolution tree structure is constructed, and convolution is performed on the corresponding child nodes CS. Coarse-scale nodes are introduced scale by scale from bottom to top, and information is exchanged between nodes using PAM. Several convolutional layers with kernel size C and step size S are sequentially applied to the embedded sequence in the time dimension to generate a sequence of length L / CS.

[0024] The pyramid attention module (PAM) is used to capture the temporal dependencies of the data after AVMD decomposition in different ranges, and the tree structure is used to perform self-attention. The features of different resolutions are extracted through inter-scale connections and intra-scale connections, and the dependencies of different scales are modeled.

[0025] Performing residual connection and layer normalization on the data processed by the pyramid attention module PAM, and scaling and translating the normalized output using learnable parameters;

[0026] The scaled and translated data is subjected to two linear transformations and one activation function in the feedforward layer. In the first linear transformation, the input is mapped to a high-dimensional space, and the nonlinear expression ability of the model is increased through the activation function. In the second linear transformation, the features of the high-dimensional space are mapped back to the original space.

[0027] The data processed by the feedforward layer is subjected to residual connection and layer normalization again, and then denormalized and input into the fully connected layer. The data is mapped to the output layer through the fully connected layer, and the prediction analysis results are obtained through the output layer.

[0028] Preferably, the processing process of the coarse-scale construction module CSCM is:

[0029]

[0030] In the formula, y norm Represents the normalized layer output data, y conv represents the output data of the convolutional layer, u represents the mean of the input data, δ 2represents the variance, ε represents a small positive number, γ controls the output scale, and β controls the output offset; y elu Represents the output of the activation function, and α is the scaling factor that controls the negative part.

[0031] Preferably, the processing process of the pyramid attention module PAM is:

[0032]

[0033] In the formula, left_side represents the left boundary, right_side represents the right boundary, p represents the current position, inner_size represents the window size, S i Indicates the starting position of the sequence, S i +L i Indicates the maximum boundary of the sequence;

[0034] The specific calculation formula for the interval boundary is:

[0035]

[0036] In the formula, p represents the current position, S i Indicates the current reference point, S i-1 Left reference point, S i+1 Right reference point, C i-1 Indicates the relationship between the position change in the current range and the reference point.

[0037] Preferably, the process of processing in the feedforward layer includes:

[0038] Use the forward propagation method forward to save a copy of the input data first, use the Gaussian error linear unit as the activation function after the first linear layer to smooth the nonlinear data, and apply dropout to discard the data; the same method is used to output the second linear layer, and a residual connection is added to the output. Finally, all processed data is normalized.

[0039] Preferably, the prediction analysis result is subjected to a Delta method to construct a confidence interval of the prediction result, including:

[0040] Determine the confidence level and use the standard deviation of the sample data to calculate the standard error and SE value respectively;

[0041] According to the confidence level and the SE value, the upper and lower bounds are calculated to obtain the confidence interval.

[0042] Preferably, obtaining the confidence interval is specifically:

[0043]

[0044] Where δ represents the batch dimension; N represents the number of samples; lower_bound is the lower bound of the calculation interval; upper_bound is the upper bound of the calculation interval, outputs is the predicted value, and std_error is the SE value.

[0045] Compared with the prior art, the present invention has the following advantages and technical effects:

[0046] The present invention improves the model training speed by using the Pyraformer network and improves the reliability of the result by using the Delta method. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0048] Figure 1 A flow chart of a method for analyzing drilling strain data based on an improved Pyraformer neural network according to an embodiment of the present invention;

[0049] Figure 2 Schematic diagram of the coarse-scale construction module CSCM and the pyramid attention module PAM structure of an embodiment of the present invention, wherein (a) is a schematic diagram of the coarse-scale construction module CSCM structure, and (b) is a schematic diagram of the pyramid attention module PAM structure;

[0050] Figure 3 A schematic diagram of the improved Pyraformer neural network structure of an embodiment of the present invention;

[0051] Figure 4 A comparison diagram of predicted data and original data of the Pyraformer neural network of an embodiment of the present invention;

[0052] Figure 5 This is a result diagram of the confidence interval output by the improved Pyraformer neural network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0054] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0055] The present invention proposes a method for analyzing drilling strain data based on an improved Pyraformer neural network, such as Figure 1 ,include:

[0056] Acquiring station surface strain data to be analyzed, and preprocessing the station surface strain data to be analyzed;

[0057] Input the preprocessed station surface strain data into the improved Pyraformer neural network for processing, and output the prediction analysis results; wherein the improved Pyraformer neural network is obtained by training and predicting the training set through stacked encoding layers and decoding layers;

[0058] The Delta method is used to construct a confidence interval of the prediction result for the prediction analysis result, and the confidence interval is output.

[0059] Specifically, this embodiment uses the adaptive variational mode decomposition (AVMD) method to decompose the borehole strain data and remove interference factors (annual trend and solid tide response); uses the coarse-scale construction module CSCM to convolve the corresponding sub-nodes, introduces coarse-scale nodes step by step from bottom to top, and uses the pyramid attention module PAM to establish a multi-resolution expression of the original data to obtain the long-term dependency between the data; for data prediction, the processed data is stacked by the pyramid attention module PAM, residual connection and layer normalization (Add&Norm) layer and feed forward layer (Feed Forward Layer), and the Delta method is used to construct a confidence interval for the output prediction data of the output layer to form an improved Pyraformer neural network architecture.

[0060] Furthermore, the station surface strain data to be analyzed are preprocessed, including:

[0061] Adaptive variational mode decomposition method AVMD is used to decompose the station surface strain data to be analyzed and calculate the initial window mean;

[0062] The station surface strain data to be analyzed is slid in windows, and the mean value of each window is judged by the initial window mean value. If there is abnormal day data, the window length is adjusted to decompose the abnormal day data separately, and VMD decomposition is performed on the adjusted window;

[0063] The trend terms and intrinsic mode function components obtained by decomposing the preprocessed station surface strain data are sorted by the adaptive variational mode decomposition method AVMD, the solid tide is decomposed, the annual trend and solid tide components of the data are removed, and the preprocessed station surface strain data are obtained.

[0064] Specifically, VMD decomposes the data into relatively stable subsequences on multiple frequency scales, and achieves efficient separation of intrinsic mode components (IMFs) by adaptively matching the optimal center frequency and bandwidth of each mode, and performs frequency domain segmentation on them, specifically:

[0065]

[0066] In the formula, u k (t) represents the kth mode obtained by decomposition; w k represents the center frequency corresponding to the kth mode; λ(t) is the Lagrange multiplier, which is used to constrain the optimization problem; f(t) represents the original signal; α represents the adjustment parameter, which is used to balance the weights of different optimization items; represents the Hilbert transform kernel function acting on the mode; represents the time differential operator; Represents the mode conversion to basis function representation based on center frequency.

[0067] Adaptive Variational Mode Decomposition (AVMD), compared with VMD, this method can construct a sliding window to judge and analyze the data, calculate the average value of each window, and determine whether there is data in the window that is higher than twice the initial mean. If not, the data is decomposed. Otherwise, the average value of each day in the window is compared and the abnormal day is decomposed separately.

[0068]

[0069] In the formula, m i represents the mean of the current window; L represents the initial window length; k, n represents the current data sequence number; m i,d represents the mean of one day in the window; D represents the step size; d represents the day in the window.

[0070] Furthermore, the initial window mean is calculated as follows:

[0071]

[0072] In the formula, m prev represents the initial window mean; L represents the initial window length; x[k] represents the current data sequence number;

[0073] The mean of each window is judged by the initial window mean, specifically:

[0074] m i >2*m prev ;

[0075]

[0076] In the formula, m iIndicates the mean of the window; m prev represents the initial window mean; L represents the initial window length; k, n represent the current data sequence number respectively; m i,d represents the mean of one day in the window; D represents the data points of each day; d represents the day number in the window.

[0077] Furthermore, the preprocessed station surface strain data are input into the improved Pyraformer neural network for processing (e.g. Figure 3 ),include:

[0078] The data after AVMD decomposition is aggregated at different scales through the coarse-scale construction module CSCM, a multi-resolution tree structure is constructed, and convolution is performed on the corresponding child nodes CS. Coarse-scale nodes are introduced scale by scale from bottom to top, and information is exchanged between nodes using PAM. Several convolutional layers with kernel size C and step size S are sequentially applied to the embedded sequence in the time dimension to generate a sequence of length L / CS.

[0079] The pyramid attention module (PAM) is used to capture the temporal dependencies of data in different ranges after AVMD decomposition, and features of different resolutions are extracted through inter-scale connections and intra-scale connections to model dependencies of different scales.

[0080] The data processed by the feedforward layer is subjected to residual connection and layer normalization again, and the data is denormalized and input into the fully connected layer;

[0081] The scaled and translated data is subjected to two linear transformations and one activation function in the feedforward layer. In the first linear transformation, the input is mapped to a high-dimensional space, and the nonlinear expression ability of the model is increased through the activation function. In the second linear transformation, the features of the high-dimensional space are mapped back to the original space.

[0082] The output of the feedforward layer is processed through the stacked pyramid attention module PAM, residual connection and layer normalization, and then the fully connected layer to obtain the regressed prediction data.

[0083] Specifically, the processing process of the coarse-scale construction module CSCM is:

[0084]

[0085] In the formula, y norm Represents the normalized layer output data, y conv represents the output data of the convolutional layer, u represents the mean of the input data, δ 2 represents the variance, ε represents a small positive number, γ controls the output scale, and β controls the output offset; y elu Represents the output of the activation function, and α is the scaling factor that controls the negative part.

[0086] The coarse-scale construction module CSCM aggregates the data decomposed by AVMD at different scales, constructs a multi-resolution tree structure, and introduces coarse-scale nodes from bottom to top by performing convolution on the corresponding child nodes CS, and then uses PAM to efficiently exchange information between nodes. Figure 2 As shown in (a), several convolutional layers with kernel size C and stride S are sequentially applied to the embedded sequence in the time dimension, resulting in a sequence of length L / CS. These fine-to-coarse sequences are cascaded and fed to PAM. To reduce the number of parameters and computations, each node is reduced through a fully connected layer before feeding the sequence to the cascaded convolutional layer and restored after all convolutions. This structure significantly reduces the number of parameters in the module and prevents overfitting.

[0087] Furthermore, the processing of the pyramid attention module PAM is as follows:

[0088]

[0089] In the formula, left_side represents the left boundary, right_side represents the right boundary, p represents the current position, inner_size represents the window size, S i Indicates the starting position of the sequence, S i +L i Indicates the maximum boundary of the sequence.

[0090] The specific calculation formula for the interval boundary is as follows:

[0091]

[0092] In the formula, p represents the current position, S i Indicates the current reference point, S i-1 Left reference point, S i+1 Right reference point, C i-1 Indicates the relationship between the position change in the current range and the reference point.

[0093] Specifically, the pyramid attention module PAM captures the temporal dependencies of different ranges of data after AVMD decomposition, such as Figure 2As shown in (b). This module uses a tree structure to perform self-attention, and extracts features of different resolutions through inter-scale connections and intra-scale connections to model dependencies at different scales. In the pyramid graph structure, the nodes at the bottom layer represent the observations at each moment, and the nodes at the upper layer extract features from the nodes at the lower layer. By connecting the nodes at each layer, the relationship between each node can be found. Since the nodes at the upper layer contain information extracted from the nodes at the lower layer, the nodes at the upper layer have already extracted and modeled information for a long period of time. Each layer only needs to consider the relationship between adjacent nodes, which reduces complexity. This pyramid graph structure is used to characterize temporal dependencies in sequences in a multi-resolution manner. From the inside to the outside, the lines pointing to themselves represent the self-attention of each node, the bidirectional lines within the layer represent the information exchange between nodes at the same scale, the bidirectional lines between layers represent the information exchange between nodes at different scales, and the unidirectional dotted line represents the maximum information propagation path required for information exchange between any two nodes:

[0094]

[0095] In the formula, Neighborhood set representing multi-scale attention; represents the local attention range of position l on the current layer (scale) S, A represents the window size of the local attention, L represents the feature map length, C s-1 represents the scale reduction factor; Represents the cross-scale context information obtained from the previous layer (scale s-1), which is only defined when scale s ≥ 2, otherwise it is represented as an empty set C represents the range of features selected from the lower scale, Represents the feature position of the previous level scale; Represents the global aggregation information obtained from higher scales.

[0096] Furthermore, the process of processing in the feedforward layer includes:

[0097] Use the forward propagation method forward to save a copy of the input data first, use the Gaussian error linear unit as the activation function after the first linear layer to smooth the nonlinear data, and apply dropout to discard the data; the same method is used to output the second linear layer, and a residual connection is added to the output. Finally, all processed data is normalized.

[0098] Specifically, the data processed by the pyramid attention module PAM undergoes residual connection and layer normalization (Add&Norm). Through residual connection, the input information can directly cross one or more layers and be added to the output of the subsequent layers, thereby alleviating the gradient vanishing and gradient exploding problems in deep networks, allowing the network to be expanded to deeper layers; layer normalization normalizes the output after residual connection, and then uses learnable parameters (such as beta and gamma) to scale and translate the normalized output. This can maintain the distribution stability of the data while retaining a certain degree of flexibility:

[0099]

[0100] After the residual connection and layer normalization (Add&Norm), the data undergoes two linear transformations and one activation function in the feed forward layer. The first linear transformation maps the input to a high-dimensional space, and then the activation function is used to increase the nonlinear expression ability of the model. Finally, the linear transformation is performed again to map the features of the high-dimensional space back to the original space. The data processed by the feed forward layer is subjected to residual connection and layer normalization again, and the data is denormalized and input into the fully connected layer.

[0101] The output of the stacked pyramid attention module PAM, residual connection and layer normalization (Add&Norm) and feed forward layer (Feed Forward Layer) is passed through the fully connected layer again to obtain the regressed prediction data.

[0102] Furthermore, the Delta method is used to construct the confidence interval of the prediction analysis results, including:

[0103] Determine the confidence level and use the standard deviation of the sample data to calculate the standard error and SE value respectively;

[0104] According to the confidence level and the SE value, the upper and lower bounds are calculated to obtain the confidence interval.

[0105] Specifically, the Delta method is a confidence interval calculation method based on the t distribution. Its basic principle is to use the difference between the sample mean and the population mean as the center of the confidence interval, and use the values ​​in the t distribution table to determine the width of the confidence interval. Compared with traditional confidence interval calculation methods, the Delta method has the advantages of simple calculation and high accuracy.

[0106] When using the Delta method, you first need to determine the confidence level, which is 95% or 99% in this embodiment; then calculate the standard error based on the standard deviation of the sample data as a basis for subsequent calculations; use the standard deviation of the sample data to calculate the standard error and calculate the SE value; calculate the upper and lower bounds based on the confidence level and SE value to obtain the confidence interval.

[0107]

[0108] Where δ represents the batch dimension; N represents the number of samples; lower_bound is the lower bound of the calculation interval; upper_bound is the upper bound of the calculation interval, outputs is the predicted value, and std_error is the SE value.

[0109] In order to more clearly express the technical solution of the present invention, the following specific embodiments are provided to introduce the solution:

[0110] Taking the borehole strain minute value data of the Liujiaxia earthquake monitoring station in Gansu Province as an example, the borehole strain data of the station was predicted and analyzed. The data was measured using a four-component borehole strain meter from September 1, 2020 to May 31, 2021.

[0111] a. Select the four-component data of the Liujiaxia earthquake monitoring station, interpolate and fill in the gaps of each component data, and then perform the transformation calculation to obtain the data s a .

[0112] b. Adaptive variational mode decomposition (AVMD) is used to analyze the data s a Decompose it, define the initial window as 7, the step size as 1, the data points per day as 1440, and calculate the initial window mean:

[0113]

[0114] In the formula, m prev represents the initial window mean; L represents the initial window length; x[k] represents the current data sequence number.

[0115] Slide the sequence data by window, and judge whether the mean of each window is twice the mean of the initial window. If it is not equal, calculate the mean of each day in the window and judge:

[0116] m i >2*m prev ;

[0117]

[0118] In the formula, m i Indicates the mean of the window; m prevrepresents the initial window mean; L represents the initial window length; k, n represents the current data sequence number; m i,d represents the mean of one day in the window; D represents the data points of each day; d represents the day number in the window.

[0119] If there are abnormal days, adjust the window length, decompose the abnormal days separately, and perform VMD decomposition on the updated window:

[0120] L' = L-(abnormal days)*D;

[0121] Where L' represents the adjusted window; L represents the window; and D represents the data points for each day.

[0122] The sequence data is decomposed into 1 trend term and 4 intrinsic mode function components. Adaptive variational mode decomposition (AVMD) decomposes the decomposed components from low frequency to high frequency, which can effectively decompose the solid tide. The annual trend and solid tide components of the data are removed and the input data is reconstructed.

[0123] c. Use token embedding to standardize the decomposed data and eliminate the scale differences between features. AVMD decomposed data is mapped to a high-dimensional space, and the temporal features of the data are encoded. Temporal embedding is used to combine temporal information, provide contextual information of the time series, and reveal the relationship between the periodicity and trend of time. Positional embedding is added to create a tensor of shape [max_len, 1], containing integers from 0-max_len-1.

[0124]

[0125] In the formula, pos represents the position of an element in the input sequence; i represents the index of the current encoding dimension; d model Represents the hidden dimension of the model.

[0126] The sine function (Sin) is used to positionally encode the data with even dimensions, and the cosine function (Cos) is used to encode the data with odd dimensions. The forward function is then used to ensure that each data can obtain the corresponding positional encoding. The positional information is added to the embedding vector and passed into the coarse-scale construction module (CSCM) together with the data.

[0127] Define a Bottleneck_Construct function to express the coarse-scale construction module (CSCM), in which the data is divided into an array list of window size 25, batch normalization is used to reduce internal covariate shift, and the ELU (Exponential Linear Unit) activation function is used to reduce the training difficulty when the output is all positive and avoid the gradient vanishing problem.

[0128]

[0129] In the formula, y norm Represents the normalized layer output data, y conv represents the output data of the convolutional layer, u represents the mean of the input data, δ 2 represents the variance, ε represents a small positive number, γ controls the output scale, and β controls the output offset; y elu Represents the output of the activation function, and α controls the scaling factor of the negative part.

[0130] The output of each convolutional layer is stored in a list, and all the outputs are concatenated together, and the dimension is upgraded through the self.up linear layer. The result of the convolution processing is concatenated with the original input, and processed through the self.norm normalization layer to obtain the final output data.

[0131] Use the get_mask function to generate the mask matrix required by the attention mechanism when processing data with multi-scale layers, inform the information flow between different levels and scales, and ensure correct attention focus and information fusion. The input data is treated as a layer of data, and then the size of each layer of the window is halved and added to the network in turn to obtain a pyramid-structured network. Through the intra-layer attention mechanism, the information between each element is extracted, and through the inter-layer attention mechanism, the positions of the elements that affect each other are calculated and the masks are recorded.

[0132]

[0133] In the formula, left_side represents the left boundary, right_side represents the right boundary, p represents the current position, inner_size represents the window size, S i Indicates the starting position of the sequence, S i +L i Indicates the maximum boundary of the sequence.

[0134] The specific calculation formula for the interval boundary is as follows:

[0135]

[0136] In the formula, p represents the current position, S iIndicates the current reference point, S i-1 Left reference point, S i+1 Right reference point, C i-1 Indicates the relationship between the position change in the current range and the reference point.

[0137] The inter-layer information is then passed and referenced layer by layer through the refer_points function. Finally, the RegularMask class is used to manage the mask matrix, returning a Boolean mask matrix to control the position attention in the attention mechanism and a list containing the sequence length of each layer.

[0138] Define a PositionwiseFeedForward class, use a two-layer feedforward neural network, process the input data bit by bit, and use a linear layer to normalize the dimension of the input data. Use the forward propagation method forward, first save a copy of the input data, use the Gaussian Error Linear Unit (GELU) as the activation function after the first linear layer to smooth the nonlinear data, and apply dropout to discard the data; the same method is used for the output of the second linear layer, and a residual connection is added after the output to reduce the gradient disappearance in the deep network. Finally, all processed data is normalized.

[0139] In this embodiment, two network layers are used, and the data of the two network layers are passed through a fully connected layer to output the predicted data. After receiving the regression prediction data, the standard deviation of the data is calculated according to 25 steps per time step to better assign data weights and ensure that the data characteristics at these moments can be reflected.

[0140]

[0141] Where δ represents the standard deviation of the predicted value at each time step calculated along the batch dimension (axis = 0); N represents the number of samples (i.e., batch size); lower_bound is the lower bound of the calculation interval; upper_bound is the upper bound of the calculation interval.

[0142] The standard error is then used to calculate the upper and lower bounds to obtain the final confidence interval. Data that exceeds this confidence interval is considered abnormal data.

[0143] d. Use the improved Pyraformer neural network to perform interval prediction, divide the data set into a training set and a test set, of which the first 80% is used as a training set and the other 20% is used as a test set; train and predict the sequence data, and compare the prediction results with the original data. Figure 4As shown in Figure 1, the method of this embodiment can effectively predict the borehole strain data, especially in the abnormal part of the data, the abnormal shape is well preserved. By constructing a confidence interval for the prediction result, the abnormal data can be well reflected, such as Figure 5 As shown in the figure, data outside the range are considered abnormal data.

[0144] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for analyzing drilling strain data based on an improved Pyraformer neural network, characterized in that: include: Acquiring station surface strain data to be analyzed, and preprocessing the station surface strain data to be analyzed; The pre-processed station surface strain data is input into the improved Pyraformer neural network for processing, and the prediction analysis results are output; wherein the improved Pyraformer neural network is obtained by training and predicting the training set through stacked encoding layers and decoding layers, and the training set is the station surface strain data set; The Delta method is used to construct a confidence interval of the prediction result for the prediction analysis result, and abnormal data is determined based on the confidence interval.

2. The analysis method according to claim 1, characterized in that Preprocessing the station surface strain data to be analyzed includes: Adaptive variational mode decomposition method AVMD is used to decompose the station surface strain data to be analyzed and calculate the initial window mean; Slide the station surface strain data to be analyzed in windows, and judge the mean of each window by the initial window mean. If there is abnormal day data, adjust the window length to decompose the abnormal day data separately, and perform VMD decomposition on the adjusted window; The trend terms and intrinsic mode function components obtained by the adaptive variational mode decomposition method AVMD are sorted, the solid tide is decomposed, the annual trend and solid tide components of the data are removed, and the preprocessed station surface strain data are obtained.

3. The analysis method according to claim 2, characterized in that The initial window mean is calculated as follows: In the formula, m prev represents the initial window mean; L represents the initial window length; x[k] represents the current data sequence number; The mean value of each window is judged by the initial window mean value, specifically: m i >2*m prev ; In the formula, m i Indicates the mean of the window; m prev represents the initial window mean; L represents the initial window length; k, n represent the current data sequence number respectively; m i,d represents the mean of one day in the window; D represents the data points of each day; d represents the day number in the window.

4. The analysis method according to claim 1, characterized in that The preprocessed station surface strain data are input into the improved Pyraformer neural network for processing, including: The data after AVMD decomposition is aggregated at different scales through the coarse-scale construction module CSCM, a multi-resolution tree structure is constructed, and convolution is performed on the corresponding child nodes CS. Coarse-scale nodes are introduced scale by scale from bottom to top, and information is exchanged between nodes using PAM. Several convolutional layers with kernel size C and step size S are sequentially applied to the embedded sequence in the time dimension to generate a sequence of length L / CS. The pyramid attention module (PAM) is used to capture the temporal dependencies of the data after AVMD decomposition in different ranges, and the tree structure is used to perform self-attention. The features of different resolutions are extracted through inter-scale connections and intra-scale connections, and the dependencies of different scales are modeled. Performing residual connection and layer normalization on the data processed by the pyramid attention module PAM, and scaling and translating the normalized output using learnable parameters; The scaled and translated data is subjected to two linear transformations and one activation function in the feedforward layer. In the first linear transformation, the input is mapped to a high-dimensional space, and the nonlinear expression ability of the model is increased through the activation function. In the second linear transformation, the features of the high-dimensional space are mapped back to the original space. The data processed by the feedforward layer is subjected to residual connection and layer normalization again, and then denormalized and input into the fully connected layer. The data is mapped to the output layer through the fully connected layer, and the prediction analysis results are obtained through the output layer.

5. The analysis method according to claim 4, characterized in that The processing process of the coarse-scale construction module CSCM is: In the formula, y norm Represents the normalized layer output data, y conv represents the output data of the convolutional layer, u represents the mean of the input data, δ 2 represents the variance, ε represents a small positive number, γ controls the output scale, and β controls the output offset; y elu Represents the output of the activation function, and α is the scaling factor that controls the negative part.

6. The analysis method according to claim 4, characterized in that The processing process of the pyramid attention module PAM is: In the formula, left_side represents the left boundary, right_side represents the right boundary, p represents the current position, inner_size represents the window size, S i Indicates the starting position of the sequence, S i +L i Indicates the maximum boundary of the sequence; The specific calculation formula for the interval boundary is: In the formula, p represents the current position, S i Indicates the current reference point, S i-1 Left reference point, S i+1 Right reference point, C i-1 Indicates the relationship between the position change in the current range and the reference point.

7. The analysis method according to claim 4, characterized in that The process of processing in the feedforward layer includes: Use the forward propagation method forward to save a copy of the input data first, use the Gaussian error linear unit as the activation function after the first linear layer to smooth the nonlinear data, and apply dropout to discard the data; the same method is used to output the second linear layer, and a residual connection is added to the output. Finally, all processed data is normalized.

8. The analysis method according to claim 1, characterized in that The Delta method is used to construct the confidence interval of the prediction results for the prediction analysis results, including: Determine the confidence level and use the standard deviation of the sample data to calculate the standard error and SE value respectively; According to the confidence level and the SE value, the upper and lower bounds are calculated to obtain the confidence interval.

9. The analysis method according to claim 8, characterized in that The confidence interval is obtained specifically as follows: lower_bound=outputs-1.95*std_error; upper_bound=outputs+1.95*std_error Where δ represents the batch dimension; N represents the number of samples; lower_bound is the lower bound of the calculation interval; upper_bound is the upper bound of the calculation interval, outputs is the predicted value, and std_error is the SE value.