A dam deformation prediction method and system
By combining K-means clustering and the HST-M model with an autoencoder dimensionality reduction method, the problem of insufficient utilization of the correlation of multiple measurement points in dam deformation prediction was solved, resulting in more stable and accurate prediction results and improving the overall risk early warning capability of dam safety monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGJIANG RIVER SCI RES INST CHANGJIANG WATER RESOURCES COMMISSION
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-24
AI Technical Summary
Existing methods for predicting dam deformation fail to effectively utilize the correlation between multiple measuring points, resulting in unstable prediction results, easy introduction of redundant information, large computational load, and inconsistencies between clustering results and actual stress distribution, affecting prediction accuracy and stability.
The K-means algorithm was used to perform cluster analysis on multiple measuring points of the dam. The optimal number of clusters was determined by combining the correlation variation method. An HST-M model was established for spatial feature fusion. The feature dimensionality was reduced by an autoencoder. Finally, a ridge regression model was used for prediction.
It improves the stability and accuracy of dam deformation prediction, can more comprehensively reflect the overall deformation behavior of the dam structure, enhances the early warning capability for potential overall risks, and maintains high performance and computational efficiency.
Smart Images

Figure CN121935893B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of dam deformation monitoring, and more specifically, to a method and system for predicting dam deformation. Background Technology
[0002] In hydraulic engineering, dam deformation is a crucial indicator for evaluating structural service performance and operational safety. Multi-point monitoring can reflect the overall deformation characteristics of the dam at different locations and times, which is of great significance for timely detection of potential safety hazards. Existing dam deformation prediction research often employs methods that only model individual monitoring points, failing to effectively utilize the correlations between different monitoring points and thus making it difficult to comprehensively reflect the overall deformation state of the dam. Other methods directly process all monitoring points uniformly, but in practical applications, this easily introduces redundant information, increases computational load, and may lead to unstable prediction results. To reduce modeling scale, some studies have introduced data grouping and spatial analysis methods; however, grouping accuracy is limited by the initial conditions of the algorithm and evaluation indicators, and clustering results may be inconsistent with the actual stress distribution, affecting prediction reliability. When fusing multi-point information for prediction, high-dimensional data often contains redundancy and noise, and the strong correlations between variables can easily lead to excessive model complexity, decreased computational efficiency, and the risk of overfitting, affecting prediction accuracy and stability. Therefore, how to fully utilize the spatial correlation information of multiple monitoring points while reducing data redundancy, minimizing noise interference, and improving prediction stability remains a pressing technical problem to be solved in the field of dam deformation monitoring and prediction. Summary of the Invention
[0003] The purpose of this application is to provide a method and system for predicting dam deformation, which can improve the predictive stability of dam deformation.
[0004] This application is implemented as follows: In a first aspect, this application provides a method for predicting dam deformation, including: The K-means algorithm was used to perform cluster analysis on multiple measuring points of the dam, merging measuring points with similar characteristics into the same group to obtain several measuring point groups; in the cluster analysis process, the correlation variation method was used to determine the optimal number of clusters. The clustered measurement point group is input into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam d、 Water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; An autoencoder is used to reduce the dimensionality of the high-dimensional feature matrix. The reduced-dimensional features are input into a ridge regression model for prediction to obtain deformation prediction results for dam measuring points.
[0005] Based on the first aspect, determining the optimal number of clusters for the K-means algorithm includes the following steps: Based on the acquired multi-point monitoring data, the correlation matrix between each monitoring point was calculated using the Pearson correlation coefficient formula, and the average intra-cluster correlation was calculated for different cluster numbers K. The calculation formula is as follows:
[0006] In the formula, The average correlation coefficient of the k-th cluster; Let be the number of samples in the k-th cluster; Let sample point i within the cluster be the cluster center. The Pearson correlation coefficient; Calculate the correlation change rate based on the average intra-cluster correlation for different K values. Its formula is:
[0007] in: The average correlation coefficient when the number of clusters is K; The number of clusters is The average correlation coefficient at that time; Select the rate of change of correlation The K value corresponding to the position that reaches the maximum is taken as the optimal number of clusters.
[0008] Furthermore, it also includes: For each K value, perform the following steps: a. Randomly select K distinct samples from all the sample data, and generate the initial cluster centers according to the following formula. :
[0009] In the formula: It is the center of the j-th cluster; Let be the number of samples in the j-th cluster; These are sample points within a cluster; Let j be the sample set of the j-th cluster; b. Calculate the Euclidean distance from each sample point to the center of each cluster, according to the formula:
[0010] Assign sample points to the nearest cluster center; c. Update the center of each cluster according to formula (3); d. Repeat steps b and c until the objective function J, i.e., the total intra-cluster error, satisfies the convergence condition or reaches the preset number of iterations. The objective function J is:
[0011] e. Calculate the average intra-cluster correlation for the clustering results based on this K value. And calculate the correlation change rate according to formula (2). ; Comparison of different K values The optimal number of clusters is determined by selecting the K value that results in the largest rate of change. K-means clustering is re-executed according to the optimal number of clusters K to obtain the optimal clustering result for multiple measurement points.
[0012] Based on the first aspect, according to the displacement of any point on the dam d、 Water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) The specific steps for establishing an HST model include: Introducing the displacement of any point on the dam d and its water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) Construct the following formula:
[0013] Water level components φ(H) The displacement of the dam body and foundation under the action of hydrostatic pressure in the reservoir consists of three parts, which are represented as follows: This refers to the displacement caused by the internal forces within the dam under hydrostatic pressure. This refers to the dam body displacement caused by the foundation displacement under the action of internal forces. The displacement of the dam body caused by the rotation of the foundation under the action of water gravity; the horizontal displacement of the arch dam due to hydrostatic pressure and water depth. and related; and The formulas for the water level components, representing the second, third, and fourth nonlinear effects of water level changes on dam deformation, are shown below:
[0014] In the formula: H The upstream water level, These are statistical coefficients; Temperature components f ( TThis describes the thermal displacement caused by temperature changes in the dam's concrete and foundation. When the temperature field inside the dam is basically stable or only observed temperature data is available, a periodic harmonic function is chosen for calculation, as shown in the following formula:
[0015] In the formula: b 1i b 2i is a statistical coefficient, and t is the cumulative number of days since the initial monitoring date; Time-sensitive quantity φ(θ) It describes the displacements caused by the changes in the material and structural properties of the dam's concrete and foundation over time; φ(θ) We can simulate this using a linear combination of a linear function and a logarithmic function, as shown in the following formula:
[0016] Where C1 and C2 are statistical coefficients, and θ = t / 100; In summary, the HST model formula is as follows:
[0017] Based on the first aspect, the specific steps for establishing the HST-M model by introducing the spatial coordinates (x, y, z) of the measurement points on the basis of the HST model include: To effectively characterize the overall deformation properties of the dam, an HST-M model was established by introducing the spatial coordinates (x, y, z) of the measuring points:
[0018] Where H, T, and θ are hydraulic, temperature, and time-dependent factors, respectively; (x, y, z) are spatial coordinate variables. f (x, y, z) is the displacement field of the dam body and dam foundation due to external loads at a given time; The spatial coordinates (x, y, z) of each measuring point are expanded using a power series method to form spatial feature fusion components. f (x, y, z):
[0019] in For statistical coefficients, k, l, m, and n represent the exponents of water level, x-coordinate, y-coordinate, and z-coordinate, respectively, and t represents the cumulative number of days relative to the reference time.
[0020] Based on the first aspect, the specific steps for using an autoencoder to reduce the dimensionality of a high-dimensional feature matrix include: The autoencoder structure is divided into an encoder and a decoder, both of which are two fully connected neural networks with a ReLU activation function after each layer. The high-dimensional feature matrix is compressed into a low-dimensional latent space by an encoder to obtain a low-dimensional feature representation. The structure formula is shown below:
[0021] Where X is the original input data, To reconstruct the data, Z is the latent representation, called the bottleneck layer. f_encoder and f_decoder are the mapping functions of the encoder and decoder, respectively. θ_encoder is the encoder parameter, including weights and biases, which determines how the encoder processes the input data. During training, the parameter θ_encoder is continuously adjusted to minimize the reconstruction error. θ_decoder is the decoder parameter, including the decoder weights and biases, which determines how the decoder transforms the latent representation Z back into the reconstruction of the original data. The original feature matrix is reconstructed based on the low-dimensional feature representation by the decoder, with mean squared error (MSE) as the loss function, and the Adam algorithm is used for optimization.
[0022] Based on the first aspect, the steps of inputting the dimensionality-reduced low-dimensional features into the ridge regression model for prediction to obtain the deformation prediction results of the dam measuring points specifically include: The ridge regression model uses L2 regularization to reduce the impact of multicollinearity on the prediction results. The function formula is shown below:
[0023] in, It represents the L2 norm, λ>0 indicates the regularization strength, w is the model weight, and n is the number of samples in the dataset. Let i be the true value of the i-th sample. Let p be the predicted value of the i-th sample, p be the number of feature dimensions of the input data, and J be the label of the feature, i.e. the j-th feature. The dimensionality-reduced feature data is input into the ridge regression model for training, and a prediction model is generated. The deformation trend of the dam measuring points is predicted based on the prediction model.
[0024] Secondly, this application also provides a dam deformation prediction system, comprising: The clustering module is configured to use the K-means algorithm to perform cluster analysis on multiple measurement points of the dam, merging measurement points with similar characteristics into the same group to obtain several measurement point groups; in the cluster analysis process, the correlation variation method is used to determine the optimal number of clusters; The fusion module is configured to input the clustered measurement point group into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam... d、 Water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; The dimension reduction module is configured to perform feature dimension reduction on the high-dimensional feature matrix using an autoencoder. The prediction module is configured to input the dimensionality-reduced low-dimensional features into the ridge regression model for prediction, in order to obtain the deformation prediction results of the dam measuring points.
[0025] Thirdly, this application also provides an electronic device, comprising: Memory, used to store one or more programs; processor; The above method is implemented when the one or more programs are executed by the processor.
[0026] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0027] Compared with the prior art, this application has at least the following advantages or beneficial effects: 1. After K-means clustering and partitioning, the high-dimensional feature matrix is obtained by inputting it into the HST-M model for spatial feature fusion. This effectively captures and utilizes the complex nonlinear spatial interactions between multiple measurement points, fundamentally overcoming the inherent defects of "information fragmentation and one-sided prediction" in single-point prediction.
[0028] 2. The introduction of an autoencoder for intelligent extraction and dimensionality reduction of high-dimensional features rich in spatial correlation information significantly improves model accuracy, especially in spatially sensitive areas. This not only leads to higher prediction accuracy, but more importantly, it integrates overall spatial state information, enabling the prediction results to more comprehensively and realistically reflect the overall deformation behavior and safety status of the dam structure, significantly enhancing the early warning capability for potential overall risks.
[0029] 3. Furthermore, the collaborative design of each component ensures that the entire framework maintains high performance while also possessing good robustness and computational efficiency. This research provides a new approach to advancing dam safety monitoring from an "isolated single-point" model to "intelligent collaboration." Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart of an embodiment of a dam deformation prediction method according to this application; Figure 2 This is a flowchart of another embodiment of a dam deformation prediction method according to this application; Figure 3 This is a heat map showing the correlation between water levels at multiple measuring points in one embodiment of a dam deformation prediction method according to this application. Figure 4 This is a schematic diagram illustrating the elbow method for determining the optimal cluster number in one embodiment of a dam deformation prediction method according to this application. Figure 5 This is a schematic diagram illustrating the use of the profile coefficient method to determine the optimal cluster number in one embodiment of a dam deformation prediction method according to this application. Figure 6 This is a schematic diagram illustrating the use of the correlation variation method to determine the optimal cluster number in one embodiment of a dam deformation prediction method according to this application. Figure 7 This is the first cluster of prediction results and residual diagram in an embodiment of a dam deformation prediction method of this application; Figure 8 This is the fifth cluster of prediction results and residual diagram in an embodiment of a dam deformation prediction method of this application; Figure 9 This is a schematic diagram of the structure of an embodiment of a dam deformation prediction system according to this application; Figure 10 This is a schematic diagram of the structure of a storage medium according to this application.
[0032] icon: 1. Clustering module; 2. Fusion module; 3. Dimensionality reduction module; 4. Prediction module; 5. Processor; 6. Memory; 7. Communication interface. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0034] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the various embodiments and features described below can be combined with each other. Example
[0035] This application provides a method and system for predicting dam deformation, which can improve the predictive stability of dam deformation.
[0036] Please refer to Figure 1-2 This method for predicting dam deformation includes the following steps: S1: The K-means algorithm is used to perform cluster analysis on multiple measuring points of the dam, and measuring points with similar characteristics are merged into the same group to obtain several measuring point groups; in the cluster analysis process, the correlation change method is used to determine the optimal number of clusters. Specifically, the K-means clustering algorithm can divide data into several disjoint groups, with highly similar data in the same group and significant differences between different groups. This embodiment uses cluster analysis to identify the changing trends of dam displacement and divides spatially similar measuring points into several groups. Further prediction is then performed on measuring points in different groups, improving prediction accuracy and computational efficiency while preserving key data. This clustering-based prediction method not only reduces multicollinearity among variables but also fully considers the connections and heterogeneity between measuring points, significantly improving monitoring accuracy and the ability to identify spatial deformation characteristics. Furthermore, K-means is a region-based clustering method that aims to divide the dataset into K clusters, making samples within clusters as similar as possible while samples between clusters are as different as possible. However, a key issue with the K-means algorithm is determining the optimal number of clusters. Different numbers of clusters directly affect the quality of the clustering results, and it is difficult to choose a suitable value without prior guidance. Therefore, this embodiment selects one method from the elbow method, the silhouette coefficient method, and the correlation variation method to determine the optimal number of clusters.
[0037] To further understand the technical solution of this application, the following engineering case cluster analysis is presented in this embodiment; Taking a hydropower station's water-retaining structure (a concrete double-curvature arch dam) as an example, this power station primarily serves power generation, while also considering flood control, navigation, and promoting local economic and social development. The dam site controls a drainage area of 406,100 km², with a multi-year average flow of 3830 m³ / s and a multi-year average runoff of 120.7 billion m³. The total installed capacity is 10200 MW, with an average annual power generation of 38.91 billion kWh. The reservoir's normal water level is 975.00 m, dead water level is 945.00 m, and check flood level is 986.17 m. The total reservoir capacity is 7.408 billion m³, with a regulating capacity of 3.02 billion m³, providing seasonal regulation capabilities. The dam crest elevation is 988.00 m, the maximum dam height is 270 m, and the dam crest arc length is 326.95 m. The spillway structures mainly consist of 5 surface outlets, 6 central outlets, and 3 flood discharge tunnels on the left bank. This water conservancy project consists of water-retaining structures, spillway structures, and water diversion and power generation systems on both banks. To understand the deformation of the dam foundation and the foundation at different depths, two vertical lines were installed at the 8th dam section at the arch crown, one inverted vertical line was installed at the 4th dam section at the left quarter of the arch crown, and two inverted vertical lines were installed at the 12th dam section at the right quarter of the arch crown. Each vertical line in sections 4 and 12 was divided into four segments, with four measuring points in each segment. The vertical line in section 8 was divided into five segments, with five measuring points in this segment. This study selected five inverted vertical lines (IP) and thirteen vertical lines (PL) from sections 4, 8, and 12. Details of the multiple measuring point selection are shown in Table 1.
[0038] Table 1. Selection of Multiple Measurement Points
[0039] This embodiment focuses on 18 monitoring points of the dam, using a dataset from June 2019 to April 2024, with daily monitoring by the instrument. The displacement of most monitoring points exhibited some fluctuation over time, but no obvious long-term increasing or decreasing trend was observed. Monitoring indicates that the dam's displacement is closely related to water level changes, especially the upstream water level, which affects the stress distribution and overall stability of the dam. When the water level rises or falls locally, the displacement of the monitoring points also shows similar fluctuations, demonstrating a correlation between water level changes and displacement. To further analyze the specific relationship between water level changes and the displacement of each monitoring point, correlation analysis of the displacement sequence was incorporated, which helps to reveal the driving role of water level in the evolution of multi-point displacement. Figure 3 The figure shows a heatmap of the correlation coefficient matrix of 18 horizontal displacement sequences measured by upright and inverted plumb bobs at a dam in this embodiment. The figure reveals strong local correlations among the displacement sequences, as indicated by the correlation coefficients. Most of them are above 0.8.
[0040] To determine the optimal number of clusters, this embodiment combines the elbow method, the profile coefficient method, and the correlation variation method to discuss the clustering of the above-mentioned multi-measurement points.
[0041] K-means clustering based on the elbow method (EM-KM) calculates the sum of squares within each cluster (SSE) for different numbers of clusters K and plots the curve of SSE as a function of K, as shown in Figure 4. According to the principle of the elbow method, the optimal number of clusters is the K value corresponding to the inflection point of the curve (i.e., the "elbow"). SSE gradually decreases as the number of clusters K increases. When K=3, the rate of decrease in SSE slows down significantly, forming a distinct "elbow".
[0042] Depend on Figure 4 It can be seen that 3 is the optimal cluster value, that is, the number of measurement points in each cluster is 9, 46 and 3. The specific cluster situation is shown in Table 2.
[0043] Table 2: EM-KM clustering results
[0044] K-means clustering based on the silhouette coefficient method (SM-KM) comprehensively considers intra-cluster compactness and inter-cluster separation. It calculates the silhouette coefficient for each data point and uses the average value as a measure of overall clustering quality. The value ranges from [-1, 1], with larger values indicating better clustering results. Figure 5 As shown, the profile coefficient reaches its maximum value when the K value is 4.
[0045] Depend on Figure 5 It can be seen that the multiple measurement points are divided into 4 clusters, with the number of measurement points in each cluster being 9, 2, 3, and 4, respectively. The specific situation within each cluster is shown in Table 3.
[0046] Table 3: SM-KM Clustering Results
[0047] K-means clustering based on correlation variation (CA-KM) calculates the degree of correlation variation under different K values based on the correlation matrix between data points, and seeks the optimal number of clusters to maximize the correlation of data points within clusters and minimize the correlation between clusters. For example... Figure 6 As shown, the mean correlation value increases with the increase of K, and a significant inflection point appears at K=5, after which the trend of change tends to level off.
[0048] Therefore, the optimal number of clusters for the correlation change method is 5, that is, the number of measurement points in each cluster is 5, 2, 2, 4, and 5. The specific cluster situation is shown in Table 4.
[0049] Table 4: CA-KM Clustering Results
[0050] Since the above three methods only evaluate clustering performance from a single dimension, they may not fully reflect the clustering quality under different numbers of clusters. Therefore, this embodiment combines the average correlation coefficient, the standard deviation of the correlation coefficient, and the average within-cluster variance to comprehensively evaluate the clustering results, improving the objectivity and stability of cluster number selection. A higher average correlation coefficient indicates a high degree of similarity among data points within a cluster; a smaller standard deviation of the correlation coefficient indicates a more uniform distribution of correlation among data points within a cluster; and a smaller average within-cluster variance indicates a higher degree of aggregation of data points within the same cluster. The optimal K-value methods are analyzed based on the evaluation criteria, as shown in Table 5.
[0051] Table 5: Clustering Evaluation Indicators
[0052] As shown in Table 5, the K-means clustering method based on correlation variation (CA-KM) identifies a more refined grouping structure of measurement points compared to the other two models. This model has the highest average correlation coefficient, indicating the strongest linear correlation between measurement points within each cluster and reasonable grouping. The lowest average intra-cluster variance indicates that the model can more tightly cluster similar measurement points, reducing intra-group dispersion, which is corroborated by the high correlation coefficient. The intra-cluster correlation stability of CA-KM is also superior to other models, as evidenced by its lower standard deviation of the correlation coefficient. Combined analysis shows that CA-KM outperforms the other two models in terms of data similarity, uniformity, and compactness, making it suitable for dam deformation prediction scenarios with high requirements for measurement point grouping accuracy. Therefore, this embodiment adopts CA-KM.
[0053] In some embodiments of the present invention, determining the optimal number of clusters for the K-means algorithm includes the following steps: Based on the acquired multi-point monitoring data, the correlation matrix between each monitoring point was calculated using the Pearson correlation coefficient formula, and the average intra-cluster correlation was calculated for different cluster numbers K. The calculation formula is as follows:
[0054] In the formula, The average correlation coefficient of the k-th cluster; Let be the number of samples in the k-th cluster; Let sample point i within the cluster be the cluster center. The Pearson correlation coefficient; Calculate the correlation change rate based on the average intra-cluster correlation for different K values. The formula is:
[0055] in: The average correlation coefficient of the k-th cluster; The number of clusters is The average correlation coefficient at that time; Select the rate of change of correlation The K value corresponding to the position that reaches the maximum is taken as the optimal number of clusters.
[0056] Furthermore, it also includes: For each K value, perform the following steps: a. Randomly select K distinct samples from all the sample data, and generate the initial cluster centers according to the following formula. :
[0057] In the formula: It is the center of the j-th cluster; Let be the number of samples in the j-th cluster; These are sample points within a cluster; Let j be the sample set of the j-th cluster; b. Calculate the Euclidean distance from each sample point to the center of each cluster, according to the formula:
[0058] Assign sample points to the nearest cluster center; c. Update the center of each cluster according to formula (3); d. Repeat steps b and c until the objective function J, i.e., the total intra-cluster error, satisfies the convergence condition or reaches the preset number of iterations. The objective function J is:
[0059] e. Calculate the average intra-cluster correlation for the clustering results based on this K value. And calculate the correlation change rate according to formula (2). ; Comparison of different K values The optimal number of clusters is determined by selecting the K value that results in the largest rate of change. K-means clustering is re-executed according to the optimal number of clusters K to obtain the optimal clustering result for multiple measurement points.
[0060] S2: The clustered measurement point group is input into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam d、 Water level components φ(H) Temperature components f ( T) and time-sensitive components φ(θ) Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; Based on the displacement of any point on the dam d、 Water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) The specific steps for establishing an HST model include: Introducing the displacement of any point on the dam d and its water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) Construct the following formula:
[0061] Water level components φ(H) The displacement of the dam body and foundation under the action of hydrostatic pressure in the reservoir consists of three parts, which are represented as follows: This refers to the displacement caused by the internal forces within the dam under hydrostatic pressure. This refers to the dam body displacement caused by the foundation displacement under the action of internal forces. The displacement of the dam body caused by the rotation of the foundation under the action of water gravity; the horizontal displacement of the arch dam due to hydrostatic pressure and water depth. and related; and The formulas for the water level components, representing the second, third, and fourth nonlinear effects of water level changes on dam deformation, are shown below:
[0062] In the formula: H The upstream water level, These are statistical coefficients; Temperature components f ( T This describes the thermal displacement caused by temperature changes in the dam's concrete and foundation. When the temperature field inside the dam is basically stable or only observed temperature data is available, a periodic harmonic function is chosen for calculation, as shown in the following formula:
[0063] In the formula: b 1i b 2i is a statistical coefficient, and t is the cumulative number of days since the initial monitoring date; Time-sensitive quantity φ(θ) It describes the displacements caused by the changes in the material and structural properties of the dam's concrete and foundation over time; φ(θ) We can simulate this using a linear combination of a linear function and a logarithmic function, as shown in the following formula:
[0064] Where C1 and C2 are statistical coefficients, and θ = t / 100; In summary, the HST model formula is as follows:
[0065] Furthermore, the steps to establish the HST-M model by introducing the spatial coordinates (x, y, z) of the measurement points based on the HST model specifically include: To effectively characterize the overall deformation properties of the dam, an HST-M model was established by introducing the spatial coordinates (x, y, z) of the measuring points:
[0066] Where H, T, and θ are hydraulic, temperature, and time-dependent factors, respectively; (x, y, z) are spatial coordinate variables. f (x, y, z) is the displacement field of the dam body and dam foundation due to external loads at a given time; The spatial coordinates (x, y, z) of each measuring point are expanded using a power series method to form spatial feature fusion components. f (x, y, z):
[0067] in For statistical coefficients, k, l, m, and n represent the exponents of water level, x-coordinate, y-coordinate, and z-coordinate, respectively, and t represents the cumulative number of days relative to the reference time.
[0068] In this step, by constructing an HST-M model, the clustered measurement point groups are input into the pre-established HST-M model to fuse spatial features and obtain a high-dimensional feature matrix, which can improve the accuracy and comprehensiveness of multi-measurement point prediction.
[0069] S3: Use an autoencoder to reduce the dimensionality of the high-dimensional feature matrix; Specifically, while the high-dimensional feature matrix output by the HST-M model can improve the accuracy and comprehensiveness of multi-point prediction, the high dimensionality of the feature matrix increases computational complexity. Therefore, this step employs an autoencoder to extract and reduce the dimensionality of the input variables, effectively capturing the spatial correlations in dam deformation. The autoencoder effectively removes redundant information, extracts key features, and completes low-dimensional representation of high-dimensional data. The encoder-decoder utilizes nonlinear mapping to improve its efficiency in extracting factors influencing complex environments, thereby enhancing the model's generalization performance and effectively mitigating overfitting.
[0070] Specifically, the autoencoder structure is divided into an encoder and a decoder, both of which are two fully connected neural networks, with the ReLU activation function used after each layer; The high-dimensional feature matrix is compressed into a low-dimensional latent space by an encoder to obtain a low-dimensional feature representation. The structure formula is shown below:
[0071] Where X is the original input data, To reconstruct the data, Z is the latent representation, called the bottleneck layer. fencoder and fdecoder are the mapping functions of the encoder and decoder, respectively. θ_encoder is the encoder's parameter, including weights and biases, which determines how the encoder processes the input data. During training, the parameter θ_encoder is continuously adjusted to minimize the reconstruction error. θ_decoder is the decoder's parameter, including the decoder's weights and biases, which determines how the decoder transforms the latent representation Z back into the reconstruction of the original data. The original feature matrix is reconstructed by the decoder based on the low-dimensional feature representation, using mean squared error (MSE) as the loss function, and optimized using the Adam algorithm.
[0072] S4: Input the dimensionality-reduced low-dimensional features into the ridge regression model for prediction to obtain the deformation prediction results of the dam measuring points.
[0073] The dimensionality-reduced low-dimensional features are input into the ridge regression model. When dealing with multivariate data, ridge regression effectively solves the multicollinearity problem through L2 regularization, further improving the model's stability and ensuring the reliability of its prediction results.
[0074] Specifically, the ridge regression model uses L2 regularization to reduce the impact of multicollinearity on the prediction results. The function formula is shown below:
[0075] in, It represents the L2 norm, λ>0 indicates the regularization strength, w is the model weight, and n is the number of samples in the dataset. Let i be the true value of the i-th sample. Let p be the predicted value of the i-th sample, p be the number of feature dimensions of the input data, and J be the label of the feature, i.e., the j-th feature. The dimensionality-reduced feature data is input into the ridge regression model for training, and a prediction model is generated. The deformation trend of the dam measuring points is predicted based on the prediction model.
[0076] In summary, this dam deformation prediction method can effectively balance feature selection and model complexity, thus maintaining high accuracy and stability in the prediction results.
[0077] To further understand this technical solution, the correlation variation method was used to determine the optimal number of clusters (K=5) for partitioning, and predictive analysis was performed on the data of the measuring points within each cluster. Referring to Table 4, clusters 1 and 5 contain 5 measuring points, covering a wide range. Compared to other clusters (such as clusters 2 and 3, which only have 2-3 measuring points), their statistical representativeness is stronger, and they can more comprehensively reflect the deformation characteristics of the region. Simultaneously, the time series data of the measuring points within clusters 1 and 5 have fewer missing data points, their displacement change trends are relatively consistent, and there is a strong deformation regularity among the measuring points. Therefore, this embodiment selects 10 measuring points—PL01DB12, PL02DB08, PL02DB12, PL03DB04, PL03DB08 in cluster 1 and PL03DB12, PL04DB04, PL04DB08, PL04DB12, PL05DB08 in cluster 5—to construct a predictive model adapted to the characteristics of each measuring point, in order to more accurately depict its displacement trend.
[0078] The data selected for this study includes monitoring points in clusters 1 and 5. Cluster 1 was monitored from June 2019 to April 2024, and cluster 5 from April 2020 to April 2024, with a monitoring frequency of once per day. During automated monitoring, multiple data values may appear at the same time point. The averaging method is typically used to ensure data consistency and reliability. Because displacement data prediction and analysis involves a large amount of deformation data, water level data, etc., from multiple monitoring points, the data may be contaminated to varying degrees, leading to various errors and anomalies. Some monitoring data may also contain redundancy or missing data, which can interfere with model training and testing. To address the problem of missing continuous time series data, linear interpolation is used for data completion. This method effectively fills in missing data, providing complete and accurate interpolated data for subsequent model training and analysis, improving data availability and the stability of analysis results.
[0079] After determining the measurement points for clusters 1 and 5 in this data selection, the dam deformation prediction method of this application was used for prediction. The data was divided chronologically, with July 2023 as the dividing line. Data before July 2023 was used for training, and data after July 2023 was used for prediction and validation. To evaluate the prediction effect of the dam deformation prediction method of this application, this embodiment compares the residuals of the model prediction results with the actual displacement data and analyzes the error distribution at different measurement points. The residual comparison diagram of the prediction results and the measured displacement data is shown below. Figure 7 , Figure 8 As shown. By Figure 7 , Figure 8It can be seen that the predicted results (training set and test set) are highly consistent with the trends of the actual observed data, indicating that the prediction method can capture the changing patterns of the displacement of the measuring points well. Within the training data segment, the prediction method shows good trend fitting; in the test set segment, the predicted values maintain a good match with the measured values, with small errors, indicating good generalization ability of the prediction method on new data. The residuals are generally small and evenly distributed, and the residuals (error components) of the 10 measuring points are generally small, with the errors of the prediction method all within a reasonable range. There are no obvious systematic biases in the overall prediction, and the model does not exhibit serious underfitting or overfitting problems. The results show that this prediction method can not only accurately fit the trend but also effectively control the error, resulting in strong generalization ability.
[0080] Please refer to Figure 9 This application also provides a dam deformation prediction system, comprising: Clustering module 1 is configured to use the K-means algorithm to perform cluster analysis on multiple measuring points of the dam, merging measuring points with similar characteristics into the same group to obtain several measuring point groups; wherein, during the cluster analysis process, the correlation variation method is used to determine the optimal number of clusters; Fusion module 2 is configured to input the clustered measurement point group into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam d、 Water level components φ(H) Temperature components f ( T ) and time-sensitive components φ(θ) Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; Dimensionality reduction module 3 is configured to perform feature dimensionality reduction on the high-dimensional feature matrix using an autoencoder; Prediction module 4 is configured to input the dimensionality-reduced low-dimensional features into the ridge regression model for prediction in order to obtain the deformation prediction results of the dam measuring points.
[0081] For a detailed implementation method of this dam deformation prediction method, please refer to the above-mentioned detailed implementation method of the dam deformation prediction system, which will not be elaborated further here.
[0082] Please refer to Figure 10 This application also provides an electronic device, including: Memory 6 is used to store one or more programs; Processor 5; Processor 5 and memory 6 are connected via communication interface 7; When the one or more programs are executed by the processor 5, all or part of the above methods are implemented.
[0083] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor 5, implements all or part of the methods described above.
[0084] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this application. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for predicting dam deformation, characterized in that, include: The K-means algorithm was used to perform cluster analysis on multiple measuring points of the dam, merging measuring points with similar characteristics into the same group to obtain several measuring point groups; in the cluster analysis process, the correlation variation method was used to determine the optimal number of clusters. The clustered measurement point group is input into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam δ、 Water level components Temperature components and time-sensitive components Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; An autoencoder is used to reduce the dimensionality of the high-dimensional feature matrix. The reduced low-dimensional features are input into the ridge regression model for prediction to obtain the deformation prediction results of the dam measuring points. Determining the optimal number of clusters for the K-means algorithm includes the following steps: Based on the acquired multi-point monitoring data, the correlation matrix between each monitoring point was calculated using the Pearson correlation coefficient formula, and different cluster numbers were calculated. Average intra-cluster correlation The calculation formula is as follows: (1) In the formula, For the first The average correlation coefficient of each cluster; For the first The number of samples in each cluster; Let sample point i within the cluster be the cluster center. The Pearson correlation coefficient; According to different The average intra-cluster correlation of the values is used to calculate the rate of change of correlation. Its formula is: in: The number of clusters is The average correlation coefficient at that time; The number of clusters is The average correlation coefficient at that time; Select the rate of change of correlation The position corresponding to the maximum The value is used as the optimal number of clusters.
2. The dam deformation prediction method according to claim 1, characterized in that, Also includes: For each To determine the value, perform the following steps: a. Randomly select from all sample data For each distinct sample, the initial cluster centers are generated according to the following formula. : In the formula: It is the center of the j-th cluster; Let be the number of samples in the j-th cluster; These are sample points within a cluster; Let j be the sample set of the j-th cluster; b. Calculate the Euclidean distance from each sample point to the center of each cluster, according to the formula: Assign sample points to the nearest cluster center; c. Update the center of each cluster according to formula (3); d. Repeat steps b and c until the objective function J, i.e., the total intra-cluster error, satisfies the convergence condition or reaches the preset number of iterations. The objective function J is: e. Regarding The clustering results of the values are used to calculate the average correlation within each cluster. And calculate the correlation change rate according to formula (2). Compare the differences Value Select the one with the largest rate of change The value is used as the optimal number of clusters; According to the optimal number of clusters Re-execute K-means clustering to obtain the optimal clustering results for multiple measurement points.
3. The dam deformation prediction method according to claim 1, characterized in that, Based on the displacement of any point on the dam δ、 Water level components Temperature components and time-sensitive components The specific steps for establishing an HST model include: Introducing the displacement of any point on the dam δ and its water level components Temperature components and time-sensitive components Construct the following formula: ; Water level components The displacement of the dam body and foundation under the action of hydrostatic pressure in the reservoir consists of three parts, which are represented as follows: This refers to the displacement caused by the internal forces within the dam under hydrostatic pressure. This refers to the dam body displacement caused by the foundation displacement under the action of internal forces. The displacement of the dam body caused by the rotation of the foundation under the action of water gravity; the horizontal displacement of the arch dam due to hydrostatic pressure and water depth. and related, and These represent the second, third, and fourth nonlinear effects of water level changes on dam deformation, respectively; the formulas for the water level components are shown below: In the formula: H The upstream water level, These are statistical coefficients; Temperature components This describes the thermal displacement caused by temperature changes in the dam's concrete and foundation. When the temperature field inside the dam is basically stable or only observed temperature data is available, a periodic harmonic function is chosen for calculation, as shown in the following formula: ; In the formula: b 1i b 2i is a statistical coefficient, and t is the cumulative number of days since the initial monitoring date; Time-sensitive quantity It describes the displacements caused by the changes in the material and structural properties of the dam's concrete and foundation over time; We can simulate this using a linear combination of a linear function and a logarithmic function, as shown in the following formula: ; Where C1 and C2 are statistical coefficients. θ =t / 100; In summary, the HST model formula is as follows: (10).
4. The dam deformation prediction method according to claim 3, characterized in that, The specific steps for establishing the HST-M model by introducing the spatial coordinates (x, y, z) of the measurement points based on the HST model include: To effectively characterize the overall deformation properties of the dam, an HST-M model was established by introducing the spatial coordinates (x, y, z) of the measuring points: Where H, T, and θ are water level, temperature, and time factor, respectively; (x, y, z) are spatial coordinate variables. It is the displacement field of the dam body and dam foundation due to external loads at a given time; The spatial coordinates (x, y, z) of each measuring point are expanded using a power series method to form spatial feature fusion components. : in, For statistical coefficients, l, m, and n represent the exponents of water level, x-coordinate, y-coordinate, and z-coordinate, respectively, and t represents the cumulative number of days relative to the reference time.
5. The dam deformation prediction method according to claim 1, characterized in that, The specific steps for using an autoencoder to reduce the dimensionality of a high-dimensional feature matrix include: The autoencoder structure is divided into an encoder and a decoder, both of which are two fully connected neural networks with a ReLU activation function after each layer. The high-dimensional feature matrix is compressed into a low-dimensional latent space by an encoder to obtain a low-dimensional feature representation. The structure formula is shown below: Where X is the original input data, To reconstruct the data, Z is an implicit representation called the bottleneck layer, and fencoder and fdecoder are the mapping functions of the encoder and decoder; The original feature matrix is reconstructed based on the low-dimensional feature representation by the decoder, with mean squared error (MSE) as the loss function, and the Adam algorithm is used for optimization.
6. The dam deformation prediction method according to claim 1, characterized in that, The specific steps involved in inputting the dimensionality-reduced low-dimensional features into a ridge regression model for prediction to obtain deformation prediction results for dam monitoring points include: The ridge regression model uses L2 regularization to reduce the impact of multicollinearity on the prediction results. The function formula is shown below: in, It represents the L2 norm, λ>0 indicates the regularization strength, w is the model weight, and n is the number of samples in the dataset. Let i be the true value of the i-th sample. Let p be the predicted value of the i-th sample, p be the number of feature dimensions of the input data, and j be the label of the feature, i.e., the j-th feature. The dimensionality-reduced feature data is input into the ridge regression model for training and a prediction model is generated. The deformation trend of the dam measuring points is predicted based on the prediction model.
7. A dam deformation prediction system, characterized in that, include: The clustering module is configured to use the K-means algorithm to perform cluster analysis on multiple measurement points of the dam, merging measurement points with similar characteristics into the same group to obtain several measurement point groups; in the cluster analysis process, the correlation variation method is used to determine the optimal number of clusters; The fusion module is configured to input the clustered measurement point group into a pre-established HST-M model for spatial feature fusion to obtain a high-dimensional feature matrix; the method for establishing the HST-M model includes: based on the displacement of any point on the dam... δ、 Water level components Temperature components and time-sensitive components Establish an HST model, and on the basis of the HST model, introduce the spatial coordinates (x, y, z) of the measurement points to establish an HST-M model; The dimension reduction module is configured to perform feature dimension reduction on the high-dimensional feature matrix using an autoencoder. The prediction module is configured to input the dimensionality-reduced low-dimensional features into the ridge regression model for prediction, in order to obtain the deformation prediction results of the dam measuring points. Determining the optimal number of clusters for the K-means algorithm includes: Based on the acquired multi-point monitoring data, the correlation matrix between each monitoring point was calculated using the Pearson correlation coefficient formula, and different cluster numbers were calculated. Average intra-cluster correlation The calculation formula is as follows: In the formula, For the first The average correlation coefficient of each cluster; For the first The number of samples in each cluster; Let sample point i within the cluster be the cluster center. The Pearson correlation coefficient; According to different The average intra-cluster correlation of the values is used to calculate the rate of change of correlation. Its formula is: in: The number of clusters is The average correlation coefficient at that time; The number of clusters is The average correlation coefficient at that time; Select the rate of change of correlation The position corresponding to the maximum The value is used as the optimal number of clusters.
8. An electronic device, characterized in that, include: Memory, used to store one or more programs; processor; When the one or more programs are executed by the processor, the method as described in any one of claims 1-6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.