Supporting structure stress state monitoring method based on artificial intelligence

By using a deep temporal neural network architecture and an adaptive wavelet attention feature mapping layer, the problems of outlier interference and weak deformation identification in stress state monitoring of support structures are solved, and efficient and accurate stress state monitoring is achieved.

CN120524256BActive Publication Date: 2025-12-09SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510708226.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-12-09
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing methods for monitoring the stress state of support structures rely on global statistics, which are easily affected by outliers. Fixed wavelet basis functions are difficult to cover the deformation characteristics of different structural parts and under multiple working conditions. Conventional optimization methods lack the ability to identify weak deformations in strain data, making it difficult to ensure that the training process proceeds smoothly and conforms to the laws of engineering physics.

Method used

A deep temporal neural network architecture is adopted, which combines an adaptive wavelet attention feature mapping layer, a temporal gated convolution module, and a dynamic feature importance reweighting layer. Through quantile-aware momentum adaptive learning rate adjustment and loss function optimization, accurate monitoring of stress in the support structure is achieved.

Benefits of technology

It significantly improves the model's robustness to local anomalies, enhances its sensitivity to local trend changes in strain data, improves the accuracy and stability of stress state monitoring, and ensures that the model output trend is consistent with the structural stress evolution law.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524256B_ABST
    Figure CN120524256B_ABST
Patent Text Reader

Abstract

The application relates to a supporting structure stress state monitoring method based on artificial intelligence and belongs to the technical field of artificial intelligence and data processing. The method comprises the following steps: acquiring and labeling strain data of a supporting structure; after removing abnormal values, normalizing multi-sensor data to generate normalized strain sequences; constructing a state monitoring model, adopting a deep time sequence neural network architecture, including an input layer, an adaptive wavelet attention feature mapping layer, a time domain gated convolution module, a global maximum pooling layer, a dynamic feature importance reweighting layer and a full connection classification layer; inputting normalized data to train the model; optimizing a loss function through a quantile interval adaptive learning rate and a momentum update strategy; after processing real-time monitoring data, inputting the data into the trained model according to a time window slice, outputting four types of probabilities and taking the maximum value as a predicted state; if a plurality of windows are in early warning and danger in succession, terminal alarm is triggered. The application can improve the accuracy of supporting structure stress state monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence and data processing, and particularly relates to a supporting structure stress state monitoring method based on artificial intelligence. BACKGROUND

[0002] The supporting structure bears time-varying load and uncertain disturbance in a complex geological environment, and slight deformation of the structure often indicates potential safety hazards. If timely identification and early warning of the stress state cannot be achieved, serious structural damage or even disasters may occur. Therefore, real-time acquisition of structural deformation information by strain sensors arranged at key positions and effective evaluation of the structure state by intelligent means are important means to ensure the safe operation of the project. The existing monitoring and early warning methods have the following problems: the conventional Min-Max or Z-Score normalization method is used to process the monitoring data, which relies on global statistics and is easily disturbed by abnormal values, resulting in deviation of the normalized results from the actual physical trend, especially when there are mutations or non-stationary sections in the monitoring time series data, the stress state evolution cannot be accurately expressed; the fixed wavelet basis function is difficult to cover the deformation characteristics of different structure parts and multiple working conditions, resulting in inefficient capture of key state information and inability to distinguish between trend strain changes and transient disturbances; the existing initialization strategy lacks the ability to identify weak deformation in strain data, and shallow features are easily degraded into noise responses, which further affects the convergence path and accuracy performance of the entire network; the conventional methods such as SGD or Adam use fixed decay or exponential averaging, which lacks response mechanism for sudden fluctuations and local state switching in strain monitoring tasks, making it difficult to ensure smooth progress of the training process and adhere to the engineering physical law. SUMMARY

[0003] The application provides a supporting structure stress state monitoring method based on artificial intelligence to solve the above problems.

[0004] To achieve the above purpose, the application realizes the following technical solutions:

[0005] The application provides a supporting structure stress state monitoring method based on artificial intelligence, which includes the following steps:

[0006] S1. Obtain strain data of the supporting structure, and label the strain data. The label categories include normal state, early warning state, dangerous state, and other state.

[0007] S2. Preprocess the labeled strain data to remove abnormal values, and normalize the preprocessed strain data to obtain a multi-sensor normalized strain sequence.

[0008] S3. Construct a state monitoring model adopting a deep time sequence neural network architecture, including an input layer, an adaptive wavelet attention feature mapping layer, a time domain gated convolution module, a global maximum pooling layer, a dynamic feature importance reweighting layer, and a fully connected classification layer; input the normalized strain sequence of the multiple sensors into the state monitoring model to train the model;

[0009] S4. In the training process, the model is optimized by a loss function, an adaptive learning rate adjustment based on quantile interval is adopted, and the parameters of the network are updated by a momentum adaptive update strategy with quantile perception to obtain a trained model;

[0010] S5. The monitored strain data is cut into segments according to a time window after being processed in step S2 to form an input sequence; the input sequence is input into the trained state monitoring model to output four types of state probabilities, and the category with the maximum probability is taken as the final prediction category; when the prediction categories of consecutive multiple time windows are “warning state” or “dangerous state”, an early warning signal is pushed to a patrol terminal.

[0011] Further, step S1 specifically comprises:

[0012] The strain data of the support structure is obtained by strain sensors arranged at key positions; the key positions include anchor rods, supports, and surrounding rock surfaces; the strain sensors include fiber Bragg gratings, strain gauges, and MEMS strain acquisition devices; a data acquisition system acquires strain data output by the strain sensors in a timed manner through a multi-channel acquisition terminal, and the storage format adopts a CSV or HDF5 structured file;

[0013] The labeling of the strain data depends on construction logs, engineering monitoring reports, and on-site manual inspection records, and the labeling method adopts expert judgment based on strain change amplitude, growth rate, and historical trend in a time period, and is recorded on the corresponding time period.

[0014] Further, step S2 specifically comprises:

[0015] In the preprocessing process, the strain data is processed by a sliding window median filter and an outlier detection mechanism to identify and eliminate isolated outliers or missing data points; a time series interpolation algorithm is used to fill in small missing data, the filling method is linear interpolation, and drift correction is performed on the sensor data to eliminate zero point offset caused by system error;

[0016] In the normalization process, for the original strain value of each sensor at time point t, a sliding time window is constructed with the time point as the center, the 0.1 quantile and the 0.9 quantile of all strain data in the window are calculated, the current strain value is subtracted from the 0.1 quantile as the numerator, and the difference between the 0.9 quantile and the 0.1 quantile is added to a constant As the denominator, the original strain is mapped to the interval [0, 1] through the scaling transformation, and the normalized strain value is obtained.

[0017] Further, the adaptive wavelet attention feature mapping layer in step S3 is specifically:

[0018] The normalized strain data is generated into a plurality of adaptive wavelet basis functions, each adaptive wavelet basis function controls the waveform position and width by adjusting the time displacement parameter and the scale parameter; for each adaptive wavelet basis function, the matching degree of the input data in the time domain is calculated, and the attention weight is calculated by querying the interaction of the vector transformation matrix, the key vector transformation matrix and the value vector transformation matrix; the adaptive wavelet basis function and the corresponding attention weight are multiplied point by point and summed to obtain the feature mapping result of fusing multi-scale features and dynamic importance allocation, which is expressed as follows:

[0019] ,

[0020] Among them, indicates the feature mapping function; indicates the normalized strain value of the i th sensor at the t th moment; indicates the normalized strain value of the i th sensor at the t th moment; indicates the preset number of wavelet basis functions; indicates a positive integer; indicates point-by-point multiplication; indicates the i th adaptive wavelet basis function; indicates the time domain attention weight. Further, the time domain gated convolution module in step S3 is specifically:

[0021] In the time domain gated convolution module, one-dimensional convolution operations of input features are respectively performed on the gating convolution kernel and the feature convolution kernel, the feature convolution result is compressed into a gating coefficient in the interval [0, 1] through the Sigmoid function, and the gating coefficient and the multi-scale gating convolution result are multiplied element by element to obtain the feature vector of the i th layer output; the gating convolution kernel adopts three gating convolution sub-kernels with window lengths of 3, 5 and 7; after the short, medium and long term time domain dependence features are extracted by the three gating convolution sub-kernels with window lengths of 3, 5 and 7, the multi-scale gating convolution result is obtained.

[0022]

[0023] ​​​Further, the dynamic feature importance reweighting in step S3 is specifically: the features output by the time domain gated convolution module are compressed in time dimension by global max pooling, and then a weight coefficient vector of each sensor channel is generated through a fully connected layer; the features output by the time domain gated convolution module are multiplied with the weight coefficient vector channel by channel to realize feature reweighting based on sensor importance adaptation.

[0024] Further, in step S3, a first-order difference of each sensor at adjacent time points is calculated for the normalized strain data to form a strain gradient matrix, and then spectral clustering analysis is performed on the strain gradient matrix to obtain a plurality of cluster centers, each cluster center vector is taken as a base value of a convolution kernel weight, and a normal distribution disturbance generated by covariance matrix decomposition of corresponding cluster samples is combined to complete initialization of network weights.

[0025] Further, the loss function in step S4 specifically includes:

[0026] On the basis of the cross-entropy loss, an L2 norm penalty term of the difference of the first-order derivative of the model output with respect to time and the real label is added, the difference approximation derivative of the output value and the label value at adjacent time steps is calculated, the consistency of the change trend of the model output and the physical process is constrained, and the formula is expressed as follows:

[0027] ,

[0028] Wherein, represents the total loss; represents the cross-entropy loss; represents the loss balance coefficient; represents the total time step; represents the L2 norm; represents the real label vector at the t time; represents the real label vector at the t-1 time; represents the derivative of the model output with respect to time; represents the derivative of the real label.

[0029] Further, the adaptive learning rate adjustment based on quantile interval in step S4 specifically includes:

[0030] A dynamic learning rate adjustment method is adopted, the 0.25th quantile, the 0.5th quantile and the 0.75th quantile of the loss value of the current training batch are calculated, the difference between the 0.75th quantile and the 0.25th quantile is divided by the 0.5th quantile to obtain a quantile interval ratio, and the learning rate of the next iteration is dynamically adjusted according to the product of the ratio and a preset decay coefficient.

[0031] Further, in step S4, the parameters of the network are updated by using the quantile-aware momentum adaptive update strategy, the gradient matrix of the current batch loss function with respect to the parameters is calculated to obtain quantile statistics, a quantile weighted momentum term is defined, and the training parameters are updated by combining the dynamic learning rate and the quantile weighted momentum term.

[0032] The present application has the advantages of:

[0033] The dynamic quantile normalization method proposed by the present application constructs a normalization ratio by using the 0.1th and 0.9th quantiles in a sliding window, significantly improves the robustness to local outliers, and enhances the sensitivity of the model to local trend changes of the strain data, thereby breaking through the problem of distribution deviation of traditional Min-Max or Z-Score under strong disturbance data; the present application combines a learnable adaptive wavelet basis function, obtains local feature positions by clustering, and constructs a basis function center and scale according to the local feature positions, and realizes collaborative modeling of high-frequency mutations and low-frequency trends by combining an attention mechanism for weighted mapping, thereby enhancing the state discrimination ability; in view of the problem that small gradient changes in strain data are difficult to effectively capture, the present application uses a first-order difference to construct a gradient matrix, extracts representative gradient features by spectral clustering, and uses the gradient features for initial weight construction of the neural network, thereby improving the initial training effect and reducing the probability of local optimal trap; by using a label and output time derivative difference penalty, the consistency of the model output trend and the structural stress evolution law is ensured, and the learning rate and gradient momentum are dynamically adjusted by combining quantile statistics, thereby solving the poor adaptability of traditional optimizers to non-stationary sequences, significantly improving the stability and convergence efficiency of the training process, and thereby improving the accuracy of the support structure stress state monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0034] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, illustrate the application, and are used together with the embodiments of the application to explain the application, and do not constitute a limitation on the application.

[0035] Figure 1 The step flowchart of the method of the present application is shown in Figure 1.

[0036] Figure 2 The classification performance comparison of different normalization methods is shown in Figure 2.

[0037] Figure 3 The effect comparison of different normalization methods is shown in Figure 3.

[0038] Figure 4 The comparison of the original signal and the frequency spectrum features after the gated convolution processing is shown in Figure 4.

[0039] Figure 5 The loss curve comparison of different model training is shown in Figure 5. DETAILED DESCRIPTION

[0040] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0041] Embodiment 1

[0042] In this embodiment, as shown in the accompanying drawings, Figure 1 The present application provides a support structure stress state monitoring method based on artificial intelligence, and the specific steps include:

[0043] S1. Obtain the strain data of the support structure, and label the strain data. The label categories include normal state, early warning state, dangerous state and other state.

[0044] Specifically, the strain data of the support structure is obtained by strain sensors arranged at key positions; the key positions include anchor rods, supports and surrounding rock surfaces; the strain sensors include fiber Bragg gratings, strain gauges and MEMS strain acquisition devices; and the small deformation changes of the structure under different working conditions are collected in real time. The data acquisition system collects the strain data output by the strain sensors through a multi-channel acquisition terminal at regular intervals, and the storage format adopts CSV or HDF5 structured file;

[0045] The labeling of the strain data depends on the construction log, engineering monitoring report and on-site manual inspection record. The labeling method is to judge the strain change amplitude, growth rate and historical trend in a time period by experts, and record it on the corresponding time period at the same time.

[0046] S2. Preprocess the labeled strain data to eliminate outliers, and normalize the preprocessed strain data to obtain a multi-sensor normalized strain sequence.

[0047] Specifically, there may be problems such as sensor drift, sudden jump, data loss, etc. in the original strain data, so cleaning operation is needed.

[0048] In the preprocessing process, the strain data is processed by sliding window median filtering and outlier detection mechanism to identify and eliminate isolated outliers or data missing points; a time series interpolation algorithm is used to fill in small missing data, the filling method is linear interpolation, and the sensor data is drift corrected to eliminate the zero point offset caused by system error;

[0049] During the normalization process, strain data is characterized by drastic temporal fluctuations and frequent local outliers. Conventional global normalization methods such as Min-Max or Z-Score will cause overall distribution shifts due to outliers and cannot capture the dynamic changes of local time windows. For the raw strain value of each sensor at time point t, a sliding time window is constructed with that time as the center. The 0.1 quantile and 0.9 quantile of all strain data within the window are calculated. The current strain value is subtracted from the 0.1 quantile as the numerator, and the difference between the 0.9 quantile and the 0.1 quantile is added to a constant. As the denominator, this proportional transformation maps the original strain to the [0,1] interval, yielding the normalized strain value, as expressed in the formula below:

[0050] ,

[0051] in, This represents the initial strain value of the i-th sensor at time t; Indicates the first The sensor at the first The strain value after time-normalization; Indicates a time index; This represents the 0.9 quantile of the strain data within the sliding window centered at time t; This represents the 0.1 quantile of the strain data within the sliding window centered at time t; This represents a constant to prevent the denominator from being zero, for example, set as... ;

[0052] It should be explained that quantiles are positional indicators that divide data into several equal parts after sorting. They reflect the local characteristics of the data distribution. For example, the 0.1 quantile is the value that makes 10% of the data appear below it, and the 0.9 quantile is the value that makes 90% of the data appear below it. Quantiles are used to eliminate the influence of extreme values ​​and enhance the ability to depict the local trends of the data.

[0053] S3. Construct a state monitoring model, which adopts a deep temporal neural network architecture, including an input layer, an adaptive wavelet attention feature mapping layer, a temporal gated convolution module, a global max pooling layer, a dynamic feature importance reweighting layer, and a fully connected classification layer; input the normalized strain sequence of multiple sensors into the state monitoring model to train the model.

[0054] Specifically, strain data has multi-scale characteristics of low-frequency trends and high-frequency abrupt changes, and feature mapping is needed to enhance the discriminative power. However, the fixed basis functions of conventional wavelet transform are difficult to adapt to different working conditions.

[0055] The normalized strain data is generated into a plurality of adaptive wavelet basis functions, each adaptive wavelet basis function controls the waveform position and width by adjusting the time displacement parameter and the scale parameter; for each adaptive wavelet basis function, the matching degree of the adaptive wavelet basis function and the input data in the time domain is calculated, and the attention weight is calculated by querying the interaction of the key vector transformation matrix, the value vector transformation matrix and the value vector transformation matrix; the adaptive wavelet basis functions and the corresponding attention weights are multiplied point by point and summed to obtain a feature mapping result fused with multi-scale features and dynamic importance allocation, which is expressed by the following formula:

[0056] ,

[0057] wherein, represents a feature mapping function; represents the normalized strain value of the i th sensor at the t th moment; represents the normalized strain value of the i th sensor at the t th moment; represents a preset wavelet basis function number; represents a positive integer; represents point-by-point multiplication; represents the i th adaptive wavelet basis function; the calculation formula is: represents an exponential function with a natural constant as the base; represents a cosine function; represents the time displacement parameter of the i th wavelet basis function; represents the scale parameter of the i th wavelet basis function; represents the time domain attention weight, and the calculation formula is: represents a Softmax function; represents a query vector transformation matrix; represents the transpose of the key vector transformation matrix; represents the i th value vector; represents the hidden layer dimension of the attention mechanism. It should be noted that the time displacement parameter is used to control the position of the wavelet basis function on the time axis, specifically, the input signal is subjected to local maximum value detection to obtain a peak point position set in each time window, then the peak point positions are arranged in time sequence, and K initial cluster centers are generated by K-means clustering, then each wavelet basis function corresponds to a cluster center, i.e. the time coordinate of the k th cluster center, which is expressed by the following formula:

[0058]

[0059] ,​​​​​​​

[0060] wherein, denotes the kth cluster center; denotes the time coordinate of the peak point belonging to the kth cluster center; denotes the number of samples within the cluster.

[0061] It should be further noted that the scale parameter is used to control the waveform width of the wavelet basis function, and its value is determined by the time span of the samples within the corresponding cluster center. Specifically, the standard deviation of the time coordinates of all peak points within the kth cluster center is calculated, and then the standard deviation is mapped to the preset scale range through exponential transformation, which is expressed as:

[0062] ,

[0063] wherein, denotes the preset minimum scale; denotes the preset maximum scale; denotes the scaling coefficient, such as 0.5; denotes the time coordinate standard deviation of the kth cluster center.

[0064] Specifically, the conventional deep neural network adopts random initialization, which is not sensitive to the local gradient changes of the support structure strain data, and is prone to cause the shallow feature extractor to fail to capture the tiny deformation features, so that the model falls into local optimum at the initial training stage;

[0065] The first-order difference of each sensor at adjacent time points is calculated for the normalized strain data to form a strain gradient matrix, and then spectral clustering analysis is performed on the strain gradient matrix to obtain a plurality of cluster centers. Each cluster center vector is used as the base value of the convolution kernel weight, and the normal distribution disturbance generated by the covariance matrix decomposition of the corresponding cluster samples is combined to complete the initialization of the network weight. The specific steps are as follows:

[0066] (1) The first-order difference of adjacent time points is calculated along the time dimension for the normalized strain data, and a matrix composed of strain gradients of all sensors at all time points is constructed, which is represented as:

[0067] ,

[0068] wherein, denotes the matrix composed of strain gradients of all sensors at all time points; denotes the first-order difference of the normalized strain data in the time dimension; denotes the matrix construction method according to and indices;

[0069] (2) Spectral clustering is performed on to generate Cluster center , This indicates the preset number of clusters. Indicates the first Cluster center vectors, Represents a positive integer.

[0070] (3) Using the cluster center vector as the initial weight base value of the corresponding convolution kernel, and superimposing the normal distribution noise generated by the square root decomposition of the sample covariance matrix of that cluster, the network weight initialization is completed, expressed as:

[0071] ,

[0072] in, Indicates the first Initial weights for each convolutional kernel; Indicates the first Calculate the covariance matrix for cluster samples; Represents the covariance matrix The square root decomposition of ; This represents a normal distribution with a mean of 0 and a variance of 1. This represents the noise injection coefficient, for example, set to 0.05.

[0073] Specifically, conventional convolutional layers have difficulty distinguishing the contribution of features at different time scales in strain data, which can easily lead to confusion between normal conditions and over-limit warnings.

[0074] In the temporal gated convolution module, one-dimensional convolution operations are performed on the input features using both gated convolution kernels and feature convolution kernels. The feature convolution result is compressed into gate coefficients in the [0,1] interval using the Sigmoid function. These gate coefficients are then multiplied element-wise with the multi-scale gated convolution result to obtain the first... The layer outputs a feature vector; the gated convolution kernel uses three gated convolution sub-kernels with time window lengths of 3, 5, and 7 respectively, which extract short, medium, and long-term temporal dependent features and then concatenate them to obtain a multi-scale gated convolution result; the formula for the temporal gated convolution operation is expressed as follows:

[0075] ,

[0076] in, This represents a one-dimensional convolution operation; Indicates the first Feature vectors output by the layer; Indicates the first The feature vector of the layer input; This indicates a temporal gated convolution operation, which uses multi-scale convolution kernels to convolve the input features and extract short, medium, and long-term features. Indicates the gated convolution kernel; denotes a characteristic convolution kernel; denotes an element-wise product; denotes a Sigmoid function.

[0077] It should be noted that, As a gating term, it dynamically adjusts the feature transmission strength and suppresses the information flow in the noise period.

[0078] It should also be noted that the convolution kernel is designed as a multi-scale structure. By performing independent convolution operations on the input features using time window length 3, 5 and 7 gated convolution sub-kernels respectively, short, medium and long term time domain features are extracted, and the channel dimensions are spliced to form a multi-scale gated convolution result. The multi-scale gated convolution kernel is represented as:

[0079] ,

[0080] wherein, denotes a time window length of 3 gated convolution sub-kernel; denotes a time window length of 5 gated convolution sub-kernel; denotes a time window length of 7 gated convolution sub-kernel; denotes a transpose operation.

[0081] Specifically, to cope with the importance difference of sensors at different positions of the support structure, the features output by the time domain gated convolution module are compressed in time dimension by global maximum pooling, and then the weight coefficient vector of each sensor channel is generated through a fully connected layer. Then, the original features are multiplied with the coefficient vector channel by channel to realize adaptive feature reweighting based on sensor importance. The formula is as follows:

[0082] ,

[0083] ,

[0084] wherein, denotes the final feature after weighting, which is used as input to the preset Softmax function to obtain a class probability vector, and then the maximum class probability is used to determine the final classification class; denotes the original feature output by the time domain gated convolution module; denotes the weight coefficient vector, which automatically learns the contribution weight of each sensor channel and suppresses the influence of noise channels; denotes a fully connected weight matrix; denotes a global maximum pooling operation.

[0085] S4. In the training process, the model is optimized by a loss function, an adaptive learning rate adjustment based on quantile interval is adopted, and a momentum adaptive update strategy with quantile awareness is used to update the parameters of the network to obtain a trained model.

[0086] Specifically, the traditional cross-entropy loss does not consider the influence of the strain change rate on the state classification, which is easy to cause misjudgment of the gradual change process.

[0087] On the basis of the cross-entropy loss, an L2 norm penalty term of the difference of the first-order derivative of the model output and the real label with respect to time is added, the approximate derivative of the difference between the output value and the label value at adjacent time steps is calculated, and the consistency of the change trend of the model output and the physical process is constrained, which is expressed by the following formula:

[0088]

[0089] Among them, represents the total loss; represents the cross-entropy loss; represents the loss balance coefficient, for example, set to 0.3; represents the total time step; represents the L2 norm; represents the real label vector at the t time; represents the real label vector at the t-1 time; represents the derivative of the model output with respect to time, by constraining the consistency of the model output and the derivative of the real label, it is ensured that the prediction of the model on the gradual change process conforms to the physical law, and it is prevented that the model only relies on the instantaneous value classification and ignores the continuity of the time sequence change; represents the derivative of the real label, which is calculated by the difference between the labels at adjacent time steps, that is , which reflects the rate of change of the stress state in the actual project;

[0090] It should be noted that, Through the L2 norm penalty term of the time derivative, the model not only considers the instantaneous strain value, but also considers the rate of change of the stress state, avoiding the shortcomings of the conventional method that only relies on the instantaneous value and ignores the time sequence change.

[0091] Specifically, the traditional learning rate decay strategy ignores the change characteristics of the loss distribution in the training process.

[0092] A dynamic learning rate adjustment method is adopted, the 0.25th quantile, the 0.5th quantile and the 0.75th quantile of the loss value of the current training batch are calculated, the difference between the 0.75th quantile and the 0.25th quantile is divided by the 0.5th quantile as the quantile interval ratio, and the learning rate of the next iteration is dynamically adjusted according to the product of the ratio and the preset decay coefficient, which is expressed by the following formula: ​

[0093] ,

[0094] wherein, denotes the learning rate of the next iteration; denotes the learning rate of the current iteration; denotes the learning rate decay coefficient, such as being set to 0.02; denotes the 0.75 quantile of the current batch loss value, reflecting the higher level of the loss value in the current batch; denotes the 0.25 quantile of the current batch loss value, reflecting the lower level of the loss value in the current batch; denotes the 0.5 quantile of the current batch loss value, representing the central tendency of the loss distribution; denotes the set composed of the loss values of all samples in the current training batch.

[0095] It should be noted that the quantile interval ratio is set to , which is used to quantify the discrete degree of the loss. The larger the interval is, the worse the prediction stability of the model on different samples is, and the learning rate needs to be reduced to stabilize the training.

[0096] Specifically, for the gradient matrix of the current batch loss function with respect to the parameters, the quantile statistics thereof is calculated; the quantile weighted momentum term is defined; and the training parameters are updated in combination with the dynamic learning rate and the momentum term.

[0097] The traditional Adam optimizer does not consider the non-stationarity and local gradient mutation of the supporting structure strain data when updating the parameters, which is easy to cause slow convergence speed;

[0098] The application adopts a quantile-aware momentum adaptive updating strategy, and the specific steps are as follows:

[0099] First, for the gradient matrix of the current batch loss function with respect to the parameters, the quantile statistics thereof is calculated, which is denoted as:

[0100] ,

[0101] wherein, denotes the gradient matrix adjusted by the quantile statistics; denotes the quantile of the gradient matrix, wherein the first parameter is the gradient matrix, and the second parameter is the value of the quantile; denotes the 0.5 quantile of the gradient matrix; denotes the 0.75 quantile of the gradient matrix; denotes the 0.25 quantile of the gradient matrix; denotes the quantile sensitivity coefficient, such as being set to 0.2.

[0102] Further, a quantile-weighted momentum term is defined, and its updating manner is represented as:

[0103] ,

[0104] wherein, denotes the quantile-weighted momentum term of the kth iteration; denotes the quantile-weighted momentum term of the (k-1)th iteration; denotes a momentum decay rate, such as being set to 0.9.

[0105] Further, in combination with the dynamic learning rate and the momentum term, the updating of the training parameter is represented as:

[0106] ,

[0107] wherein, denotes the model parameter of the (k+1)th iteration; denotes the model parameter of the kth iteration; denotes the training parameter of the neural network; is a robustness adjustment coefficient, such as being set to 0.01; is a sign function; denotes a square root operation on the absolute value, used for suppressing the nonlinear scaling of the gradient amplitude.

[0108] When the total loss changes by less than 0.01 in consecutive 5 iteration steps, the iteration is stopped, indicating that the model converges, that is, the model training is completed.

[0109] S5. The monitored strain data is processed through step S2 and divided into segments according to a time window, such as setting 60 seconds as the time window to form an input sequence; the input sequence is input into the trained state monitoring model to output four types of state probabilities, and the category with the largest probability is taken as the final prediction category; when the prediction categories of three consecutive time windows are “warning state” or “dangerous state”, a warning signal is pushed to the inspection terminal.

[0110] Embodiment 2

[0111] In this embodiment, the influence of different normalization methods on the classification performance is compared through the box plot, verifying the superiority of the sliding window quantile normalization proposed in the present application over the traditional global normalization method, such as Figure 2As shown in the figure, the horizontal axis represents three processing methods: global minimum-maximum normalization, standardization, and the method of this invention. The vertical axis represents classification accuracy (a) and F1 score (b). Experimental results show that the box position using the method of this invention is significantly higher than that of the traditional method, and the box height is more compact with fewer outliers, indicating that the method of this invention can effectively suppress outlier interference and enhance the model's ability to capture local fluctuation features. The higher median and smaller interquartile range in the box plot demonstrate that the quantile-based dynamic normalization strategy can better preserve the local temporal characteristics of strain data, thereby improving classification stability.

[0112] like Figure 3 As shown, this paper verifies the adaptability of the proposed sliding window quantile normalization method to strain data processing compared to traditional global normalization methods. By comparing the processing effects of the proposed method with minimum-maximum normalization and standardized normalization on the same sensor data, it can be found that traditional methods produce significant numerical shifts when there are abrupt changes in the time series. Their normalization results fluctuate drastically near the abrupt change points and recover slowly. In contrast, the proposed method effectively suppresses the impact of outliers on the overall data distribution through quantile calculation within a local time window, achieving a more stable numerical mapping while maintaining the original fluctuation characteristics of the data. The time axis in the figure shows the continuous monitoring process, and the normalization value axis reflects the compression effect of different methods on the original data. The curve of the proposed method shows better anti-interference ability while retaining reasonable fluctuations, proving the necessity of local dynamic normalization for engineering monitoring data processing.

[0113] Example 3

[0114] In this embodiment, the ability of the temporal gated convolution module to capture multi-scale deformations of the support structure is verified by comparing the spectral characteristics of the original signal and the signal after gated convolution processing. Figure 4 As shown, the horizontal axis represents the signal frequency components, and the vertical axis displays the corresponding energy intensity. In the original spectrum, high-frequency noise energy is severely mixed with the effective signal, while the processed signal shows a significant increase in energy concentration at the fundamental frequency representing low-frequency structural deformation, a noticeable attenuation of high-frequency noise energy, and energy enhancement at the characteristic frequencies corresponding to sudden anomalies. Experimental results show that the gated convolution module, through multi-scale convolution kernels and dynamic weight adjustment, achieves selective enhancement of the core features of structural health status while suppressing irrelevant interference, enabling the model to accurately distinguish between gradual deformation and sudden emergencies.

[0115] Example 4

[0116] In this embodiment, training loss curves are used to compare the learning efficiency of different model architectures, verifying the synergistic effect of the adaptive wavelet attention mechanism and the temporal gated convolution structure, such as... Figure 5As shown in the figure, the horizontal axis is the training round, the vertical axis is the loss value, and the four curves respectively represent the long short-term memory network, the convolutional neural network, the self-attention mechanism network and the method of the application. The experimental results show that the initial descending rate of the curve of the application is the fastest, the fluctuation amplitude in the middle period is the smallest, and finally it is stabilized at the lowest level. It shows that the adaptive wavelet basis function can accurately match the multi-scale strain characteristics, the gating convolution module effectively filters noise interference, and the slight jitter at the end of the curve reflects the fitting of the model to the real physical law in the data, which proves that the derivative penalty term enhances the modeling ability of the model to the gradual change process.

[0117] Finally, it should be noted that: the above only for the preferred embodiments of the application, and not for the purpose of limiting the application, although the application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, it still can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.

Claims

1. An artificial intelligence-based support structure stress state monitoring method, characterized by, The method comprises the following steps: S1. Obtain strain data of the support structure, and label the strain data, with the label categories including normal state, early warning state, dangerous state, and other state; S2. Preprocess the labeled strain data to eliminate abnormal values, normalize the preprocessed strain data, and obtain a multi-sensor normalized strain sequence; S3. Construct a state monitoring model, wherein the state monitoring model adopts a deep time sequence neural network architecture, and comprises an input layer, an adaptive wavelet attention feature mapping layer, a time domain gated convolution module, a global maximum pooling layer, a dynamic feature importance reweighting layer, and a fully connected classification layer; and the multi-sensor normalized strain sequence is input into the state monitoring model to train the model; The adaptive wavelet attention feature mapping layer specifically comprises: The normalized strain data is used to generate a plurality of adaptive wavelet basis functions, each adaptive wavelet basis function controls the waveform position and width by adjusting the time shift parameter and the scale parameter; for each adaptive wavelet basis function, the matching degree of the adaptive wavelet basis function with the input data in the time domain is calculated, and the time domain attention weight is calculated through the interaction of the query vector transformation matrix, the key vector transformation matrix, and the value vector transformation matrix; the adaptive wavelet basis functions are multiplied point by point with the corresponding attention weights, and then summed to obtain a feature mapping result that fuses multi-scale features and dynamic importance allocation, which is expressed as follows: , wherein, represents a characteristic mapping function; represents a normalized strain value of the th sensor at the th time point; represents a preset wavelet basis function number; represents a positive integer; represents point-by-point multiplication; represents the th adaptive wavelet basis function; represents a time domain attention weight; S4. In the training process, the model is optimized through a loss function, an adaptive learning rate adjustment based on quantile interval is adopted, a momentum adaptive update strategy based on quantile perception is used to update the parameters of the network, and a trained model is obtained; S5. The monitored strain data is processed according to step S2, and then divided into segments according to a time window to form an input sequence; the input sequence is input into the trained state monitoring model, and four state probabilities are output; the category with the maximum probability is taken as the final prediction category; when the prediction categories of a plurality of consecutive time windows are "early warning state" or "dangerous state", an early warning signal is pushed to a patrol terminal.

2. The artificial intelligence-based support structure stress state monitoring method according to claim 1, characterized by, Step S1 specifically comprises: The strain data of the support structure is obtained by strain sensors arranged at key positions; the key positions include anchor rods, supports, and surrounding rock surfaces; the strain sensors include fiber Bragg gratings, strain gauges, and MEMS strain acquisition devices; a data acquisition system acquires strain data output by the strain sensors through a multi-channel acquisition terminal at regular time intervals, and stores the strain data in a CSV or HDF5 structured file; The labeling of the strain data relies on construction logs, engineering monitoring reports, and on-site manual inspection records; an expert judges the strain change amplitude, growth rate, and historical trend in a time period, and records the judgment on the corresponding time period.

3. The artificial intelligence-based support structure stress state monitoring method according to claim 2, characterized by, Step S2 specifically comprises: In the preprocessing process, the strain data is processed by a sliding window median filter and an abnormal value detection mechanism to identify and eliminate isolated outliers or missing data points; a time series interpolation algorithm is used to fill in small amounts of missing data, the filling method is linear interpolation, and sensor data is drift corrected to eliminate zero point deviation caused by system errors; In the normalization process, for each sensor at time point t, a sliding time window is constructed with the time point as the center, the 0.1th quantile and the 0.9th quantile of all strain data in the window are calculated, and the current strain value is subtracted from the 0.1th quantile as the numerator, and the difference between the 0.9th quantile and the 0.1th quantile is added to a constant As the denominator, the original strain is mapped to the [0, 1] interval through the proportional transformation, and the normalized strain value is obtained.

4. The artificial intelligence-based support structure stress state monitoring method according to claim 3, characterized by, The time domain gated convolution module in step S3 is specifically as follows: In the time domain gating convolution module, one-dimensional convolution operations of an input feature are respectively performed on a gating convolution kernel and a feature convolution kernel, a feature convolution result is compressed into a gating coefficient in an interval of [0, 1] through a Sigmoid function, and the gating coefficient is multiplied with a multi-scale gating convolution result element by element to obtain a feature vector output by the time domain gating convolution module. The gating convolution kernel adopts three gating convolution sub-kernels with time window lengths of 3, 5 and 7 respectively. After short-term, medium-term and long-term time domain dependence features are extracted by the three gating convolution sub-kernels with time window lengths of 3, 5 and 7 respectively, the features are spliced to obtain a multi-scale gating convolution result.

5. The method of claim 4, wherein the method further comprises: The dynamic feature importance reweighting in step S3 is specifically as follows: After the features output by the time domain gated convolution module are compressed in the time dimension through global maximum pooling, a weight coefficient vector of each sensor channel is generated through a fully connected layer; the features output by the time domain gated convolution module are multiplied with the weight coefficient vector channel by channel to realize feature reweighting based on sensor importance adaptation.

6. The artificial intelligence-based support structure stress state monitoring method according to claim 5, characterized by, In step S3, the first-order difference of each sensor at adjacent time points is calculated for the normalized strain data to form a strain gradient matrix, and then spectral clustering analysis is performed on the strain gradient matrix to obtain a plurality of cluster centers, each cluster center vector is taken as a base value of a convolution kernel weight, and a normal distribution disturbance generated by covariance matrix decomposition of corresponding cluster samples is combined to complete initialization of network weights.

7. The artificial intelligence-based support structure stress state monitoring method according to claim 6, characterized by, The loss function in step S4 specifically includes: On the basis of the cross-entropy loss, an L2 norm penalty term of the difference of the first-order derivative of the model output with respect to time and the true label is added, the approximate derivative of the difference between the output value and the label value at adjacent time steps is calculated, the consistency of the change trend of the model output and the physical process is constrained, and the formula is expressed as follows: , wherein, denotes the total loss; denotes the cross-entropy loss; denotes the loss balancing coefficient; denotes the total time step; denotes the L2 norm; denotes the real label vector at time t; denotes the real label vector at time t-1; denotes the derivative of the model output with respect to time; denotes the derivative of the real label.

8. The artificial intelligence-based support structure stress state monitoring method according to claim 7, characterized in that, The adaptive learning rate adjustment based on quantile interval in step S4 specifically includes: A dynamic learning rate adjustment method is adopted, the 0.25th quantile, the 0.5th quantile and the 0.75th quantile of the loss value of the current training batch are calculated, the difference between the 0.75th quantile and the 0.25th quantile is divided by the 0.5th quantile to obtain a quantile interval ratio, and the learning rate of the next iteration is dynamically adjusted according to the product of the ratio and a preset decay coefficient.

9. The artificial intelligence-based support structure stress state monitoring method according to claim 8, characterized by, The quantile-aware momentum adaptive update strategy is used to update the parameters of the network in step S4: quantile statistics are calculated for the gradient matrix of the loss function with respect to the parameters of the current batch; a quantile weighted momentum term is defined; and the training parameters are updated in combination with the dynamic learning rate and the quantile weighted momentum term.

Citation Information

Patent Citations

  • Urban garbage capacity sensing method

    CN118503622A

  • Dynamic health adaptive monitoring method and system using artificial intelligence

    CN119480112A