A method for constructing a stress prediction model of a high arch dam based on monitoring data

By combining Tent chaotic mapping, refraction reverse learning and adaptive t-distribution variation, the LightGBM model is optimized, and the SHAP framework is introduced, the accuracy and interpretability problems of high arch dam stress prediction are solved, and the accurate prediction of high arch dam stress and influencing factor identification are achieved.

CN115935488BActive Publication Date: 2025-08-05CHANGJIANG RIVER SCI RES INST CHANGJIANG WATER RESOURCES COMMISSION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310024859.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-08-05
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

It is difficult to establish an accurate high-arch dam stress prediction model in the prior art, especially due to the complex nonlinear characteristics of high-arch dam systems and the local minimum problems of artificial intelligence methods, resulting in insufficient prediction capabilities and poor interpretability.

Method used

The LightGBM model is optimized using Tent chaotic mapping, refraction reverse learning strategy and adaptive t-distribution variation, and combined with the SHAP framework, a high-arch dam stress prediction model is established, and the parameter optimization and model interpretability are enhanced by monitoring data.

Benefits of technology

It improves the accuracy and interpretability of stress prediction of high arch dams, can effectively identify influencing factors, and provides health monitoring and diagnostic basis for high arch dams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115935488B_ABST
    Figure CN115935488B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a high arch dam stress prediction model based on monitoring data, comprising: establishing a LightGBM high arch dam stress prediction model based on monitoring data from a high arch dam without stress gauges and a strain gauge group; employing Tent chaos mapping and a refraction reverse learning strategy to enhance the diversity and quality of the initial sparrow population, and enabling sparrows to jump out of the local optimal position based on adaptive t-distribution variation, thereby improving the global search capability of the sparrow search algorithm; optimizing and analyzing the LightGBM model using an improved sparrow search algorithm to determine the optimal hyperparameter combination, and introducing the SHAP framework for interpretable black box models to establish an interpretable high arch dam stress prediction model. The high arch dam stress prediction model of the present invention, which integrates the improved sparrow search algorithm, LightGBM, and SHAP, can accurately predict high arch dam stress, identify significant features that affect high arch dam stress, and provide a decision-making basis for high arch dam health monitoring and diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hydraulic structure safety monitoring, and in particular to a method for constructing a high arch dam stress prediction model based on monitoring data. Background Art

[0002] High arch dams are subject to significant loads during operation, potentially causing excessive stress. Excessive tensile stress can cause cracking, while excessive shear stress can lead to slippage. Therefore, excessive stress in the dam body is a primary cause of cracks and overall instability. Currently, stress prediction models for high arch dams are often developed by analyzing monitoring data from both strain gauges and strain gauges to assess the factors affecting stress and the safety status of high arch dams.

[0003] Among the many methods for analyzing dam stress and strain monitoring data and establishing prediction models, linear analysis methods such as principal component analysis or partial least squares regression are generally used to establish statistical models. The operation period of a high arch dam is affected by many factors, including upstream and downstream water levels, dam body temperature, shrinkage deformation of the mountain valley on both sides, and concrete creep. The high arch dam system is a complex nonlinear system, and conventional linear regression analysis methods cannot establish an accurate stress prediction model. Some artificial intelligence methods for solving complex nonlinear problems, such as artificial neural networks, support vector machines, and extreme learning machines, have problems such as over-learning and are prone to falling into local minima. Therefore, how to establish a high-predictive and interpretable high arch dam stress prediction model based on monitoring data is an urgent problem that needs to be solved. Summary of the Invention

[0004] In response to the problems existing in the prior art, the present invention provides a method for constructing a high arch dam stress prediction model based on monitoring data. The method is based on the analysis of monitoring data without stress gauges and strain gauge groups, integrates Tent chaos mapping, refraction reverse learning strategy and adaptive t distribution variation, proposes an improved sparrow search algorithm (ISSA) to optimize the parameters of the LightGBM high arch dam stress prediction model, and introduces the SHAP framework to enhance the interpretability of the model. The technical solution combines ISSA, LightGBM and SHAP to establish a high arch dam stress prediction model with high predictive ability and interpretability.

[0005] The technical solution adopted by the present invention to solve the technical problem is:

[0006] A method for constructing a high arch dam stress prediction model based on monitoring data comprises the following steps:

[0007] S1: Strain gauge group stress calculation: Combine the monitoring data of the high arch dam without stress gauges and the strain gauge group to calculate the uniaxial strain of the high arch dam; Based on the calculated uniaxial strain of the high arch dam, combined with the elastic modulus and creep formula obtained from concrete tests, the stress of the high arch dam is calculated using the deformation method;

[0008] S2: Establishment of LightGBM high arch dam stress prediction model: The stress of the high arch dam calculated in step S1 and the determined influencing factors of the high arch dam stress are divided into a training set and a prediction set, and a high arch dam stress prediction model is established based on the divided training set and prediction set and the LightGBM algorithm;

[0009] S3: Population initialization: Set the relevant parameters of the sparrow search algorithm and the range of LightGBM hyperparameters, and use Tent chaos mapping theory and refraction back learning to initialize the sparrow population of hyperparameters;

[0010] S4: Calculation of optimal hyperparameters: Use the LightGBM high arch dam stress prediction model established in step S2 to train the training set and calculate the fitness value of individual sparrows for the prediction set; update the positions of the discoverer, joiner, and scout, and calculate the fitness value of the sparrow population; use the adaptive t distribution to mutate the positions of all sparrows and update the position of the sparrow population; determine whether the algorithm has reached the maximum number of iterations. If not, return to calculate the fitness value of the individual sparrow; if satisfied, output the optimal hyperparameters;

[0011] S5: Establishment of an interpretable high arch dam stress prediction model: Based on the optimal hyperparameters calculated in step S4, the SHAP framework is introduced to establish an interpretable LightGBM high arch dam stress prediction model. The interpretable LightGBM high arch dam stress prediction model is used to evaluate the importance of factors affecting high arch dam stress.

[0012] Furthermore, the specific steps of calculating the stress of the strain gauge group in step S1 are as follows:

[0013] S11. Combined with the high arch dam non-stress gauge monitoring data, the non-stress gauge statistical model equation is established:

[0014] ε0=a0+a1T+a2t+a3Ln(1+t)+a4e kt (1)

[0015] Where: ε0 is the measured value of the stress-free gauge, T is the current temperature of the stress-free gauge, t is the time from the start date of the analysis, a0, a1, a2, a3, a4 are regression coefficients, and k is -0.01;

[0016] S12. Obtain the regression coefficient in step S11 by the least squares method, combine it with the monitoring data of the high arch dam strain gauge group, deduct the free volume deformation of the concrete to obtain the stress and strain ε under the external load.n The calculation formula is as follows:

[0017]

[0018] Where: is the measured value of the strain in a certain direction of the strain gauge group, is the average value of the measured temperature of the strain gauge group;

[0019] S13, stress and strain ε in any direction n n The following relationship is satisfied with the normal strain and shear strain:

[0020] ε n =ε x l 2 +ε y m 2 +ε z n 2 +γ xy lm+γ yz mn+γ zx nl (3)

[0021] Where: ε x , ε y , ε z is the normal strain in the x, y, and z coordinate axes, γ xy , γ yz , γ zx is the shear strain on the xy, yz, and zx coordinate planes, l, m, and n are the direction cosines, l = cos(n,x), m = cos(n,y), and n = cos(n,z);

[0022] Solve the following system of equations to convert the measured strains into normal strains and shear strains:

[0023] E n =Aε (4)

[0024] Where: E n is the strain of the isotropic strain gauge after deducting the free volume deformation ε n A is the coefficient matrix composed of l, m, and n, and ε is the vector of normal strain and shear strain;

[0025] S14. The calculation formula for the uniaxial strain of the strain gauge group is as follows:

[0026]

[0027] Where: ε ′ x , ε ′ y , ε z ′is the uniaxial strain in the x, y, and z coordinate axes, γ y ′ z , γ x ′ z, , γ x ′ y is the uniaxial strain on the xy, yz, and zx coordinate planes;

[0028] S15. Divide the uniaxial strain process line into a series of unequal time periods and calculate the stress increment in each time period using the following loading method:

[0029]

[0030] Where: τ i is the end age of the i-th calculation period; is the midpoint age of the i-th calculation period, ε i ′ is the uniaxial strain at the end age of the i-th calculation period; for The instantaneous elastic modulus of concrete at this moment; Expression The loading period lasts until τ i The creep degree; Indicates The unit stress of age loading lasts until τ i Total deformation The derivative of i The continuous elastic modulus at the moment; Δσ(τ i ) is τ i The stress increment at time

[0031] The stress at any moment can be obtained by superposition:

[0032]

[0033] Where: σ(τ n ) is any moment τ n arch, radial or vertical stress.

[0034] Furthermore, the specific steps for establishing the LightGBM high arch dam stress prediction model in step S2 are as follows:

[0035] S21, plotting the process lines of the high arch dam stress calculated in step S1 and the upstream water level, temperature, and valley deformation, and qualitatively analyzing the importance of the factors affecting the high arch dam stress;

[0036] S22, perform nonlinear correlation analysis by calculating the maximum information coefficient between the high arch dam stress and the influencing factors of upstream water level, temperature and valley deformation, and refer to the qualitative analysis of step S21 to delete the influencing factors with a maximum information coefficient less than 0.1, and determine the influencing factors of the high arch dam stress y x = (x1, x2, ..., x n );

[0037] S23, dividing the high arch dam stress and the stress influencing factors determined in step S22 into a training set and a prediction set;

[0038] S24. The purpose of the LightGBM algorithm is to find an approximate value of a function f(x) Minimize the expected value of a specific loss function L(y,f(x)) as follows:

[0039]

[0040] The final objective function after segmentation is:

[0041]

[0042] Where: g i and h i are the first-order and second-order gradient statistics of the loss function, λ is the coefficient of L2 regularization, I L and I R are the sample sets of the left and right branches respectively, I=I L ∪I L is the parent sample set;

[0043] S25. Based on the training set and prediction set divided in step S23 and the LightGBM algorithm determined in step S24, a high arch dam stress prediction model is established by continuing deep optimization through vertical growth tree:

[0044]

[0045] Where: x i (i=1,2,…,n) is the influencing factor of high arch dam stress, and y is the high arch dam stress.

[0046] Furthermore, the specific steps of population initialization in step S3 are as follows:

[0047] S31. Set the relevant parameters of the sparrow search algorithm, including the population size M, the number of iterations, the warning value, the ratio of discoverers, the ratio of scouts, and the range of LightGBM hyperparameters, including the number of cotyledons num_leaves, the minimum number of cotyledon data min_data_in_leaf, the maximum depth max_depth, and the learning rate learning_rate range;

[0048] S32. For the LightGBM hyperparameters set in step S31, the sparrow population is initialized using Tent mapping, and the expression of the sparrow chaotic population is obtained as follows:

[0049]

[0050] Where: x ij is the position of the i-th sparrow in the population in the j-th dimension; ub j and lb j are the minimum and maximum values of the jth dimension of the search space respectively; y j is the j-th dimension Tent mapping, and its expression is:

[0051]

[0052] Where: 0<α<1, take α=0.7; y1 is a random number between -1 and 1;

[0053] S33. Improve the individual quality of the sparrow chaotic population obtained in step S32 through refraction reverse learning. The calculation formula for the refraction reverse position of the sparrow chaotic population is as follows:

[0054]

[0055] Where: is x ij The refraction reverse position of , k is the zoom coefficient of the lens, and other parameters are shown in formula (11) in step S32;

[0056] S34, merge the sparrow chaotic population x in step S32 ij and the refraction reverse population in step S33 Sort by the rise and fall of fitness values, and select the top M sparrow individuals in fitness values as the initial population of LightGBM hyperparameters.

[0057] Furthermore, the specific steps for calculating the optimal hyperparameters in step S4 are as follows:

[0058] S41, using the LightGBM hyperparameter sparrow population determined in step S3 and the high arch dam stress prediction model established in step S2 based on LightGBM to train the training set divided in step S23, and calculating the fitness value of each sparrow with respect to the prediction set divided in step S23; the calculation formula of the fitness value is as follows:

[0059]

[0060] Where y i is the measured value, is the predicted value, is the mean, n is the number of measured values, and RMSE is the root mean square error;

[0061] S42. Sort the individual fitness values of all sparrows in the sparrow population according to the individual fitness values calculated in step S41 to find the current best and worst fitness values;

[0062] S43. Update the discoverer's location. The formula is:

[0063]

[0064] Where: represents the position of the i-th sparrow at the t+1 iteration; is a constant, indicating the maximum number of iterations; C ξ ∈(0,1) is a random number; R2∈[0,1], ST∈[0,1] represent the warning value and safety value respectively; C Q is a random number that obeys the normal distribution; L is a multi-dimensional matrix with one row and all elements are 1;

[0065] S44. Update the joiner's position. The formula is:

[0066]

[0067] Where: X P The best position of the current discoverer; X worst Indicates the current global worst position; A is a multidimensional matrix with each element being 1 or -1. + =A T (AA T );

[0068] S45. Update the scout's position. The formula is:

[0069]

[0070] Where: is the current global best position; C β is the step size control parameter; C K ∈(0,1) is a random number; f i is the current sparrow’s fitness value, f g and f w is the current best fitness value and the worst fitness value; C ε is a constant used to avoid the denominator being zero and is set to 10e-10;

[0071] S46, using formula (14) in step S41 to calculate the updated sparrow population fitness value;

[0072] S47: The process of performing t-distribution variation on the position of individual sparrows is as follows:

[0073]

[0074] Where: x i and are the positions of the i-th sparrow before and after mutation, k is the number of iterations, is a t-distribution with the number of iterations k as the degree of freedom; the probability density function under the t-distribution with k degrees of freedom is as follows:

[0075]

[0076] S48: Calculate the fitness values of all sparrows after mutation in step S47 using formula (14) in step S41; compare with the fitness values of the sparrow population determined in step S46, re-sort, and update the position of the sparrow population;

[0077] S49: Determine whether the algorithm has reached the maximum number of iterations. If not, return to step S41; if yes, output the optimal hyperparameters.

[0078] Furthermore, the specific steps for establishing the interpretable high arch dam stress prediction model in step S5 are as follows:

[0079] S51: Determine the final LightGBM high arch dam stress prediction model based on the optimal hyperparameters calculated in step S4;

[0080] S52: For the LightGBM high arch dam stress prediction model determined in step S51, a new model interpretation method SHAP using the Shapley value of game theory is introduced. The Shapley value is calculated as follows:

[0081]

[0082] Where: F represents the set of all features, S represents all feature subsets after removing the i-th feature from F, f S∪{i} (x S∪{i} ) represents the model trained with the i-th feature, f S (x S ) represents the model trained without the i-th feature, φ i represents the Shapley value of the i-th feature;

[0083] The calculation formula of the interpretable high arch dam stress prediction model is as follows:

[0084]

[0085] Where: g is the explanatory model; z ′ ∈{0,1+M , which is 1 if the feature exists, otherwise 0; M is the number of input features.

[0086] The present invention has the following beneficial effects:

[0087] (1) This paper proposes a new and improved sparrow search algorithm, which uses tent chaotic mapping and refraction reverse learning strategy to enhance the diversity and quality of the initial sparrow population, and based on adaptive t-distribution mutation, it makes sparrows jump out of the local optimal position, thereby improving the global search capability of the sparrow search algorithm.

[0088] (2) The present invention uses an improved sparrow search algorithm to optimize and analyze the LightGBM model, determine the optimal hyperparameter combination, and introduces the SHAP framework of the interpretable black box model to establish an interpretable high arch dam stress prediction model; this model can improve training efficiency and prediction accuracy, and quantitatively analyze the positive and negative influences of factors affecting the stress of high arch dams. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 This is a flow chart of one embodiment of a method for constructing a high arch dam stress prediction model based on monitoring data of the present invention;

[0090] Figure 2 This is an overall planar layout diagram of the high arch dam strain gauge group used in an embodiment of the present invention;

[0091] Figure 3 is the spatial distribution of stress in the arch direction of the downstream strain gauge group of the high arch dam according to the embodiment of the present invention;

[0092] Figure 4 This is the process line of the arch stress and upstream water level of the typical strain gauge group S314-2 downstream of the high arch dam according to the embodiment of the present invention;

[0093] Figure 5 is a process line of the average values of the arch stress and temperature measured by a typical strain gauge group S314-2 downstream of a high arch dam according to an embodiment of the present invention;

[0094] Figure 6 This is a layout diagram of the valley deformation observation line of the high arch dam according to an embodiment of the present invention;

[0095] Figure 7 This is the process line of arch stress and valley deformation of a typical strain gauge group S314-2 downstream of a high arch dam according to an embodiment of the present invention;

[0096] Figure 8 It is the optimization iterative process of the ISSA-LightGBM, SSA-LightGBM and PSO-LightGBM models in the embodiment of the present invention;

[0097] Figure 9This is a comparison chart of the predicted values and measured values of the three models in the embodiment of the present invention;

[0098] Figure 10 is the residual graph predicted by the three models in the embodiment of the present invention;

[0099] Figure 11 : This is an importance distribution diagram of each characteristic factor of dam stress of the ISSA-LightGBM model according to an embodiment of the present invention;

[0100] Figure 12 It is a distribution diagram of Shapley values of the explanation model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0101] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention.

[0102] Please refer to Figure 1 The present invention provides a method for constructing a high arch dam stress prediction model based on monitoring data, comprising the following steps:

[0103] S1: strain gauge group stress calculation;

[0104] In order to analyze and verify the effectiveness of the proposed method, the monitoring data of the Xiluodu high arch dam without stress gauge and strain gauge group were selected for analysis; the overall plane distribution of the strain gauge group of the high arch dam is shown in Figure 2. Figure 2 ;

[0105] Since the arch compressive stress in the downstream of the arch dam is relatively large, the prediction model is mainly established for the maximum arch compressive stress of the downstream strain gauge group; the spatial distribution of the arch stress of the strain gauge group downstream of the high arch dam on October 21, 2021 is shown in Figure 3 shown by Figure 3 It can be seen that the arch stress downstream of the dam is generally in a compressive state, and the high-pressure stress area is concentrated in the middle and lower part of the downstream riverbed dam section; the current maximum arch compressive stress is 10.19MPa, located in the strain gauge group S314-2 at an elevation of 372m in the 14# dam section.

[0106] S2: LightGBM high arch dam stress prediction model establishment;

[0107] According to dam engineering theory and mechanics, water pressure generates stress on the dam body. Since there is a cushion pond downstream of the high arch dam, the water level of the cushion pond remains basically unchanged. Therefore, the water pressure component only needs to consider the influence of the upstream water pressure. The process line of the arch stress and upstream water level of the typical strain gauge group S314-2 downstream of the high arch dam is shown in Figure 1. Figure 4 ;Depend on Figure 4 It can be seen that the arch stress downstream of the high arch dam is negatively correlated with the upstream water level, and the arch compressive stress tends to increase gradually.

[0108] The temperature component is the stress caused by the temperature change of the dam body; the process line of the arch stress and the average temperature value of the typical strain gauge group S314-2 downstream of the high arch dam is shown in Figure 5 ;Depend on Figure 5 It can be seen that the arch stress downstream of the high arch dam has an obvious correlation with the temperature, and the overall correlation is negative.

[0109] The valley deformation of the mountains on both sides of the high arch dam squeezes the high arch dam, causing stress changes in the dam body. Seven valley deformation observation lines (VDL01-VDR01, VDL02-VDR02, ..., VDL07-VDR07) are arranged before and after the dam, of which four are arranged in front of the dam and three are arranged behind the dam. The arrangement of the valley deformation observation lines is shown in Figure 2. Figure 6 The process line of arch stress and valley deformation of the typical strain gauge group S314-2 downstream of the high arch dam is shown in Figure 7 ;Depend on Figure 7 It can be seen that the valley deformation of the mountains on both sides has not completely converged, and the arch stress downstream of the high arch dam is positively correlated with the valley deformation overall;

[0110] The stress of high arch dams is mainly related to factors such as water pressure, temperature, valley deformation and aging. Since valley deformation is also affected by aging factors, in order to avoid duplication of independent variable factors, the stress of high arch dams only needs to consider the influence of water pressure, temperature and valley deformation. The high arch dam was fully poured to the top on March 6, 2014. The arch stress prediction model of the downstream strain gauge group S314-2 selected the monitoring data from March 11, 2014 to February 1, 2021 as the training set, with a total of 240 samples, and the monitoring data from February 15, 2021 to October 21, 2021 as the prediction set, with a total of 28 samples. The high arch dam stress prediction model established based on LightGBM is:

[0111]

[0112] Where: σ is the arch stress of the high arch dam, H is the upstream water level, T is the average value of the measured temperature of the strain gauge group, V j It is valley deformation.

[0113] S3: population initialization;

[0114] The population size of the sparrow search algorithm (SSA) is set to 20, the number of iterations is 500, the warning value is 0.6, the proportion of discoverers is 70%, and the proportion of scouts is 20%; in order to compare the superiority of the improved sparrow search algorithm (ISSA), the particle swarm algorithm (PSO) is used to optimize the LightGBM high arch dam stress prediction model established in step S2. The learning factor of PSO is set to 0.2, the acceleration constant is set to 2, and the inertia weight decreases linearly from 0.9 to 0.4; the hyperparameters of the LightGBM high arch dam stress prediction model include 4 parameters: the number of cotyledons num_leaves, the minimum number of cotyledon data min_data_in_leaf, the maximum depth max_depth and the learning rate learning_rate. The setting range is shown in Table 1; the sparrow population of the hyperparameters is initialized using Tent chaos mapping theory and refraction reverse learning.

[0115] Table 1 LightGBM model hyperparameter range

[0116] Hyperparameters scope num_leaves [2,30] min_data_in_leaf [1,10] max_depth [2,10] learning_rate [0.01,0.2]

[0117] S4: Optimal hyperparameter calculation;

[0118] Based on the LightGBM hyperparameter sparrow population determined in step S3, the LightGBM high arch dam stress prediction model established in step S2 was optimized 10 times using ISSA, SSA and PSO, and the average optimal fitness value of the 10 times was calculated;

[0119] The optimization iterative process of ISSA-LightGBM, SSA-LightGBM and PSO-LightGBM models is shown in Figure 8 ;from Figure 8 It can be seen that the ISSA-LightGBM model does not have the problem of premature convergence in terms of convergence speed and accuracy;

[0120] Comparison of the predicted values and measured values of the three models is shown in Figure 9 ;from Figure 9 It can be seen that the predicted values of the ISSA-LightGBM model are basically consistent with the measured values, which is better than SSA-LightGBM and PSO-LightGBM;

[0121] Figure 10 are the residuals predicted by the three models; Figure 10 It can be seen that compared with other prediction models, the prediction residual mean of ISSA-LightGBM is smaller and the distribution is concentrated, indicating that the prediction effect of this model is the best;

[0122] Table 2 shows the best, worst, mean, and standard deviation of the RMSE of the prediction set when each model is run 10 times. The results show that the prediction accuracy of the ISSA-LightGBM model is better than that of the other two models. The optimal values of the hyperparameters of the three LightGBM models are finally obtained and shown in Table 3.

[0123] Table 2 Comparison of RMSE of prediction sets of different models

[0124] Model Best value Worst value mean Standard deviation ISSA-LightGBM 0.077 0.077 0.077 0.000 SSA-LightGBM 0.085 0.090 0.087 1.741e-3 PSO-LightGBM 0.092 0.106 0.093 4.599e-3

[0125] Table 3 Optimal values of hyperparameters for different models

[0126]

[0127]

[0128] S5: Establishment of an interpretable stress prediction model for high arch dams;

[0129] LightGBM can output the contribution of each feature (dam stress influencing factor) to the prediction results. Figure 11 Figure 3 is the importance distribution of each characteristic factor of dam stress in the ISSA-LightGBM model. It can be seen from the figure that the water pressure component H and temperature T have the most significant influence, and the valley deformation has the least influence on the model. The high arch dam strain gauge group S314-2 is located in the middle and lower part of the downstream riverbed dam section. The arch stress in this part is less affected by the upstream water level. The buried elevation of this measuring point is higher than the water level of the cushion pond, and the stress is significantly affected by the temperature. The valley shrinkage deformation has a significant squeezing effect on the downstream dam body and has a greater impact on the arch stress in the low-elevation part downstream of the high arch dam. In summary, the valley deformation and temperature have a significant impact on the arch stress of the high arch dam strain gauge group S314-2, and the upstream water level has little effect on the arch stress. Therefore, the importance of the characteristic factors obtained by LightGBM is inconsistent with the actual situation.

[0130] Figure 12 This is the Shapley value distribution diagram of all features calculated based on the ISSA-LightGBM stress model analysis of high arch dams. Each row in the figure represents a feature, and the horizontal axis is the Shapley value. A point represents a sample. The darker the color, the larger the value of the feature itself, and the lighter the color, the smaller the value of the feature itself. It can be seen from the figure that the valley deformation V3, V1, and temperature T have the most significant impact on the model, and the water pressure component has the least impact on the model. In addition, the valley deformation is positively correlated with the dam stress, the temperature is negatively correlated with the dam stress, and the water pressure component has a small correlation with the arch stress downstream of the dam. These are consistent with the actual variation law of dam stress.

[0131] Through the implementation and experimental analysis of the present invention, the following conclusions can be finally obtained:

[0132] 1. The improved sparrow search algorithm proposed in the present invention, which integrates Tent chaotic mapping, refraction reverse learning strategy and t distribution, increases the search diversity and quality of the sparrow population and improves the global search capability.

[0133] 2. The ISSA-LightGBM dam stress prediction model proposed in this paper has higher prediction accuracy and generalization ability compared with the SSA-LightGBM and PSO-LightGBM models.

[0134] 3. The present invention uses SHAP to enhance the interpretability of the ISSA-LightGBM model, evaluate the importance of factors affecting dam stress, and identify significant features affecting dam stress.

[0135] In summary, the method for constructing a high arch dam stress prediction model based on monitoring data proposed in the present invention is effective. Compared with traditional stress prediction models, it has higher prediction accuracy and interpretability. Therefore, it is recommended for promotion and application in actual engineering monitoring.

[0136] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for constructing a high arch dam stress prediction model based on monitoring data, characterized in that: The following steps are involved: S1: Strain gauge group stress calculation: Combine the monitoring data of the high arch dam without stress gauges and the strain gauge group to calculate the uniaxial strain of the high arch dam; Based on the calculated uniaxial strain of the high arch dam, combined with the elastic modulus and creep formula obtained from concrete tests, the stress of the high arch dam is calculated using the deformation method; S2: Establishment of LightGBM high arch dam stress prediction model: The stress of the high arch dam calculated in step S1 and the determined influencing factors of the high arch dam stress are divided into a training set and a prediction set, and a high arch dam stress prediction model is established based on the divided training set and prediction set and the LightGBM algorithm; S3: Population initialization: Set the relevant parameters of the sparrow search algorithm and the range of LightGBM hyperparameters, and use Tent chaos mapping theory and refraction back learning to initialize the sparrow population of hyperparameters; S4: Calculation of optimal hyperparameters: Use the LightGBM high arch dam stress prediction model established in step S2 to train the training set and calculate the fitness value of the sparrow individual on the prediction set; Update the positions of the discoverer, joiner, and scout, and calculate the fitness value of the sparrow population; use the adaptive t-distribution to mutate the positions of all sparrows and update the position of the sparrow population; determine whether the algorithm has reached the maximum number of iterations. If not, return to calculate the fitness value of the sparrow individual; if it does, output the optimal hyperparameters; S5: Establishment of an interpretable high arch dam stress prediction model: Based on the optimal hyperparameters calculated in step S4, the SHAP framework is introduced to establish an interpretable LightGBM high arch dam stress prediction model. The interpretable LightGBM high arch dam stress prediction model is used to evaluate the importance of factors affecting high arch dam stress.

2. The method for constructing a high arch dam stress prediction model based on monitoring data according to claim 1, characterized in that: The specific steps for calculating the stress of the strain gauge group in step S1 are as follows: S11. Combined with the high arch dam non-stress gauge monitoring data, the non-stress gauge statistical model equation is established: ε0=a0+a1T+a2t+a3Ln(1+t)+a4e kt (1); Where: ε0 is the measured value of the stress-free gauge, T is the current temperature of the stress-free gauge, t is the time from the start date of the analysis, a0, a1, a2, a3, a4 are regression coefficients, and k is -0.01; S12. Obtain the regression coefficient in step S11 by the least squares method, combine it with the monitoring data of the high arch dam strain gauge group, deduct the free volume deformation of the concrete to obtain the stress and strain ε under the external load. n The calculation formula is as follows: Where: is the measured value of the strain in a certain direction of the strain gauge group, is the average value of the measured temperature of the strain gauge group; S13, stress and strain ε in any direction n n The following relationship is satisfied with the normal strain and shear strain: e n =e x l 2 +e y m 2 +e z n 2 +g xy lm+γ yz mn+g zx nl (3); Where: ε x , ε y , ε z is the normal strain in the x, y, and z coordinate axes, γ xy , γ yz , γ zx is the shear strain on the xy, yz, and zx coordinate planes, l, m, and n are the direction cosines, l = cos(n,x), m = cos(n,y), and n = cos(n,z); Solve the following system of equations to convert the measured strains into normal strains and shear strains: AND n =A ε (4); Where: E n is the strain of the isotropic strain gauge after deducting the free volume deformation ε n A is the coefficient matrix composed of l, m, and n, and ε is the vector of normal strain and shear strain; S14. The calculation formula for the uniaxial strain of the strain gauge group is as follows: Where: ε' x 、ε' y 、ε' z is the uniaxial strain in the x, y, and z coordinate axes, γ′ yz ,γ′ xz ,γ′ xy is the uniaxial strain on the xy, yz, and zx coordinate planes; S15. Divide the uniaxial strain process line into a series of unequal time periods and calculate the stress increment in each time period using the following loading method: Where: τ i is the end age of the i-th calculation period; is the midpoint age of the i-th calculation period, ε′ i is the uniaxial strain at the end age of the i-th calculation period; for The instantaneous elastic modulus of concrete at this moment; Expression The loading period lasts until τ i The creep degree; Indicates The unit stress of age loading lasts until τ i Total deformation The derivative of i The continuous elastic modulus at the moment; Δσ(τ i ) is τ i The stress increment at time The stress at any moment can be obtained by superposition: Where: σ(τ n ) is any moment τ n arch, radial or vertical stress.

3. The method for constructing a high arch dam stress prediction model based on monitoring data according to claim 1, characterized in that: The specific steps for establishing the LightGBM high arch dam stress prediction model in step S2 are as follows: S21, plotting the process lines of the high arch dam stress calculated in step S1 and the upstream water level, temperature, and valley deformation, and qualitatively analyzing the importance of the factors affecting the high arch dam stress; S22, perform nonlinear correlation analysis by calculating the maximum information coefficient between the high arch dam stress and the influencing factors of upstream water level, temperature and valley deformation, and refer to the qualitative analysis of step S21 to delete the influencing factors with a maximum information coefficient less than 0.1, and determine the influencing factors of the high arch dam stress y x = (x1, x2, ..., x n ); S23, dividing the high arch dam stress and the stress influencing factors determined in step S22 into a training set and a prediction set; S24. The purpose of the LightGBM algorithm is to find an approximate value of a function f(x) As shown below: The final objective function after segmentation is: Where: g i and h i are the first-order and second-order gradient statistics of the loss function, λ is the coefficient of L2 regularization, I L and I R are the sample sets of the left and right branches respectively, I=I L ∪I L is the parent sample set; S25. Based on the training set and prediction set divided in step S23 and the LightGBM algorithm determined in step S24, a high arch dam stress prediction model is established by continuing deep optimization through vertical growth tree: Where: x i is the influencing factor of high arch dam stress, where i = 1, 2,…, n; y is the high arch dam stress.

4. The method for constructing a high arch dam stress prediction model based on monitoring data according to claim 1, characterized in that: The specific steps of population initialization in step S3 are as follows: S31. Set the relevant parameters of the sparrow search algorithm, including the population size M, the number of iterations, the warning value, the ratio of discoverers, the ratio of scouts, and the range of LightGBM hyperparameters, including the number of cotyledons num_leaves, the minimum number of cotyledon data min_data_in_leaf, the maximum depth max_depth, and the learning rate learning_rate range; S32. For the LightGBM hyperparameters set in step S31, the sparrow population is initialized using Tent mapping, and the expression of the sparrow chaotic population is obtained as follows: x ij =(ub j -lb j )×y j +lb j (11); Where: x ij is the position of the i-th sparrow in the population in the j-th dimension; ub j and lb j are the minimum and maximum values of the jth dimension of the search space respectively; y j is the j-th dimension Tent mapping, and its expression is: Where: 0<α<1, take α=0.7; y j A random number between -1 and 1; S33. Improve the individual quality of the sparrow chaotic population obtained in step S32 through refraction reverse learning. The calculation formula for the refraction reverse position of the sparrow chaotic population is as follows: Where: is x ij The refraction reverse position of , k is the zoom coefficient of the lens, and other parameters are shown in formula (11) in step S32; S34, merge the sparrow chaotic population x in step S32 ij and the refraction reverse population in step S33 Sort by the rise and fall of fitness values, and select the top M sparrow individuals in fitness values as the initial population of LightGBM hyperparameters.

5. The method for constructing a high arch dam stress prediction model based on monitoring data according to claim 3, characterized in that: The specific steps for calculating the optimal hyperparameters in step S4 are as follows: S41, using the LightGBM hyperparameter sparrow population determined in step S3 and the high arch dam stress prediction model established in step S2 based on LightGBM to train the training set divided in step S23, and calculating the fitness value of each sparrow with respect to the prediction set divided in step S23; the calculation formula of the fitness value is as follows: Where y i is the measured value, is the predicted value, m is the number of measured values, and RMSE is the root mean square error; S42. Sort the individual fitness values of all sparrows in the sparrow population according to the individual fitness values calculated in step S41 to find the current best and worst fitness values; S43. Update the discoverer's location. The formula is: Where: represents the position of the i-th sparrow at the t+1 iteration; is a constant, indicating the maximum number of iterations; C ξ ∈(0,1) is a random number; R2∈[0,1], ST∈[0,1] represent the warning value and safety value respectively; C Q is a random number that obeys the normal distribution; L is a multi-dimensional matrix with one row and all elements are 1; S44. Update the joiner's position. The formula is: Where: X P The best position of the current discoverer; X worst Indicates the current global worst position; A is a multidimensional matrix with each element being 1 or -1. + =A T (AA T ); S45. Update the scout's position. The formula is: Where: is the current global best position; C β is the step size control parameter; C K ∈(0,1) is a random number; f i is the current sparrow’s fitness value, f g and f w is the current best fitness value and the worst fitness value; C ε is a constant used to avoid the denominator being zero and is set to 10e-10; S46, using formula (14) in step S41 to calculate the updated sparrow population fitness value; S47: The process of performing t-distribution variation on the position of individual sparrows is as follows: Where: x i and are the positions of the i-th sparrow before and after mutation, k is the number of iterations, is a t-distribution with the number of iterations k as the degree of freedom; the probability density function under the t-distribution with k degrees of freedom is as follows: S48: Calculate the fitness values of all sparrows after mutation in step S47 using formula (14) in step S41; compare with the fitness values of the sparrow population determined in step S46, re-sort, and update the position of the sparrow population; S49: Determine whether the algorithm has reached the maximum number of iterations. If not, return to step S41; if yes, output the optimal hyperparameters.

6. The method for constructing a high arch dam stress prediction model based on monitoring data according to claim 1, characterized in that: The specific steps for establishing the interpretable high arch dam stress prediction model in step S5 are as follows: S51: Determine the final LightGBM high arch dam stress prediction model based on the optimal hyperparameters calculated in step S4; S52: For the LightGBM high arch dam stress prediction model determined in step S51, a model interpretation method SHAP using the Shapley value of game theory is introduced. The Shapley value is calculated as follows: Where: F represents the set of all features, S represents all feature subsets after removing the i-th feature from F, f S∪{i} (x S∪{i} ) represents the model trained with the i-th feature, f S (x S ) represents the model trained without the i-th feature, φ i represents the Shapley value of the i-th feature; The calculation formula of the interpretable high arch dam stress prediction model is as follows: Where: g is the explanation model; z′∈{0,1} M , which is 1 if the feature exists, otherwise 0; M is the number of input features.

Citation Information

Patent Citations

  • Arch dam safety early warning method under influence of valley amplitude deformation

    CN114372393A

  • Dam state prediction method and system based on data assimilation

    CN114547951A