Method for determining dam deformation safety monitoring indexes based on factor screening

By constructing a method based on factor screening and deep autoregression model, the problem of insufficient data samples in the formulation of dam safety monitoring indicators was solved, and the accurate and reliable formulation of dam deformation safety monitoring indicators was achieved, thereby improving the applicability of the project and its safety assurance capabilities.

CN122490476APending Publication Date: 2026-07-31SHAANXI HUANGHE GUXIAN TECH INNOVATION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI HUANGHE GUXIAN TECH INNOVATION CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for formulating dam safety monitoring indicators suffer from insufficient reliability and difficulty in reflecting the combined effects of randomness and ambiguity during dam operation when data samples are insufficient. Traditional methods are unable to achieve accurate and reliable safety monitoring.

Method used

By acquiring dam deformation monitoring data and its influencing factors, an initial factor set is constructed. Correlation analysis is performed to screen out influencing factors with correlation thresholds. Kernel linear discriminant analysis is used for dimensionality reduction. A deep autoregressive model based on quantile regression improvement is constructed. The model parameters are optimized using the walrus optimization algorithm to generate the probability prediction interval of dam deformation in order to formulate safety monitoring indicators.

Benefits of technology

This approach improves the accuracy and reliability of dam deformation safety monitoring indicators even with insufficient data samples, better reflects the uncertainty of deformation, and significantly enhances the applicability and safety assurance capabilities of the project.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490476A_ABST
    Figure CN122490476A_ABST
Patent Text Reader

Abstract

This application discloses a method for formulating dam deformation safety monitoring indicators based on factor screening, relating to the field of dam safety monitoring technology. The method first acquires dam deformation monitoring data and its influencing factors to construct an initial factor set; it then eliminates weakly correlated factors through correlation screening, followed by dimensionality reduction of the screened influencing factors using kernel linear discriminant analysis to extract principal component factors; it constructs a deep autoregressive model based on quantile regression improvement, and uses the walrus optimization algorithm to globally optimize the model parameters, obtaining an optimized probability prediction model; finally, it uses this model to generate probability prediction intervals for dam deformation, formulates safety monitoring indicators, and verifies and evaluates these indicators by formulating comparative indicators using the confidence interval method. This application overcomes the shortcomings of traditional methods, such as high sample dependence, poor fitting of nonlinear data, and insufficient indicator reliability, achieving accurate formulation of dam deformation safety monitoring indicators and providing reliable support for dam safe operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of dam safety monitoring technology, and in particular to a method for formulating dam deformation safety monitoring indicators based on factor screening. Background Technology

[0002] In response to the new situation and demands for the safe and long-term operation of dams, it is crucial to accurately formulate dam deformation safety monitoring indicators, assess the dam's ability to withstand loads, and provide important scientific basis for taking preventive measures in advance under the most unfavorable load combinations, effectively controlling and evaluating dam safety. However, in complex reservoir environments, the actual deformation monitoring time series obtained are generally non-stationary, noisy, and complex data influenced by multiple factors such as water level and temperature. These data are characterized by a lack of extreme load samples in the initial stages of operation and the evolution of data distribution characteristics over time.

[0003] With the development of technology, current methods for formulating dam safety monitoring indicators are mainly divided into two categories: mathematical statistics and structural analysis. The former mainly includes confidence interval estimation and low-probability methods; the latter is mainly based on deterministic models using the finite element method. Mathematical statistics, with its advantages of simplicity, ease of implementation, and wide engineering application, has achieved fundamental application results in determining warning values ​​for effect sizes based on monitoring data.

[0004] However, traditional monitoring models have some shortcomings. For example, the typical low-probability method in mathematical statistics relies excessively on the sample size and deterministic probability distribution of the monitoring data. Its main limitation is that in actual engineering projects, it is often difficult to obtain a sufficient number of sample data under adverse load combinations, which limits the reliability of the proposed results. On the other hand, the confidence interval method is limited by the time span of the prediction model and requires periodic updates to ensure accuracy. In this context, relying solely on traditional methods fails to fully reflect the combined effects of randomness and fuzziness during dam operation. Furthermore, although structural analysis methods have clear physical concepts, they require high precision in material parameters, are computationally complex, and are difficult to apply in practice. Summary of the Invention

[0005] The purpose of this application is to provide a method for formulating dam deformation safety monitoring indicators based on factor screening, which effectively overcomes the problem of insufficient reliability of monitoring indicators caused by insufficient data samples, and provides more decision support for dynamic assessment of dam deformation safety.

[0006] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for formulating dam deformation safety monitoring indicators based on factor screening, comprising: acquiring dam deformation monitoring data and its corresponding influencing factors, and constructing an initial factor set; performing correlation analysis on each influencing factor in the initial factor set, and screening influencing factors that meet the correlation threshold according to a preset correlation threshold; performing dimensionality reduction processing on the screened influencing factors using kernel linear discriminant analysis to extract principal component factors; constructing a deep autoregressive model based on quantile regression improvement based on the extracted principal component factors, and optimizing the parameters of the deep autoregressive model; generating a probability prediction interval for dam deformation using the optimized deep autoregressive model, and formulating dam deformation safety monitoring indicators based on the probability prediction interval.

[0007] Optionally, the step of performing correlation analysis on each influencing factor in the initial factor set and filtering out influencing factors that meet the correlation threshold according to a preset correlation threshold includes: The Spearman rank correlation coefficient method was used to calculate the rank correlation coefficient between each of the influencing factors in the initial factor set and the dam deformation monitoring data. Based on the magnitude of the rank correlation coefficient, each of the influencing factors is divided into several different correlation levels; Based on the preset relevance threshold, influence factors corresponding to the relevance levels that do not reach the relevance threshold are removed, and influence factors that meet the relevance threshold are obtained.

[0008] Optionally, kernel linear discriminant analysis is used to reduce the dimensionality of the screened influencing factors and extract principal component factors, including: The selected impact factors are used as input, and the Gaussian kernel function is used to map the selected impact factors to a high-dimensional feature space. In the high-dimensional feature space, the optimal projection direction is found by maximizing the ratio of inter-class divergence to intra-class cohesion. The optimal projection direction is the projection direction that maximizes the inter-class separation and intra-class cohesion of the mapped samples in the low-dimensional space. The selected influence factors are projected into a low-dimensional space according to the optimal projection direction to obtain the dimension-reduced influence factors. Plot the cumulative contribution rate curve, and extract the principal component factors from the dimensionality-reduced influence factors whose cumulative contribution rate reaches the preset cumulative contribution rate threshold.

[0009] Optionally, the step of constructing a deep autoregressive model based on quantile regression improvement based on the extracted principal component factors, and optimizing the parameters of the deep autoregressive model, includes: Using the extracted principal component factors as input variables, a deep autoregressive model based on a recurrent neural network is constructed. The deep autoregressive model is used to describe the probability distribution of dam deformation monitoring data over time. The output of the deep autoregressive model is replaced by the distribution parameters with the quantile prediction values ​​corresponding to multiple preset quantiles, wherein the preset quantiles include upper quantiles and lower quantiles; The quantile loss function is used as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function. The model parameters are learned by minimizing the quantile loss function.

[0010] Optionally, the step of using the quantile loss function as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function, and learning the model parameters by minimizing the quantile loss function, includes: learning model parameters using the loss function. During the training process, the input of the deep autoregressive model at each time step includes the hidden layer output of the previous time step, the dam deformation monitoring data of the previous time step, and the covariates of the current time step. The hidden layer output of the current time step is obtained through the calculation of the recurrent neural network, and then the quantile loss function value of the current time step is calculated based on the hidden layer output of the current time step. By minimizing the sum of the quantile loss function values ​​of all time steps, the optimal parameters of the deep autoregressive model are learned.

[0011] Optionally, the parameter optimization of the deep autoregressive model using the walrus optimization algorithm includes: Initialize the population and determine the position of each walrus individual in the population. The position of each walrus individual represents a set of parameters to be optimized in the depth autoregressive model. During the foraging phase, each walrus individual updates its own position under the guidance of the current best individual in order to search for a better combination of parameters; During the migration phase, each walrus individual randomly selects another individual in the population and decides to move closer to or further away from that individual based on the fitness of the selected individual in order to expand the search range; During the predator evasion phase, each walrus individual conducts a local search within a small area around its current location to fine-tune its parameter combinations; The foraging phase, migration phase, and predator evasion phase are executed iteratively until the termination condition is met, and the location of the optimal walrus individual is output. The location of the optimal walrus individual is the optimal parameter combination of the depth autoregressive model.

[0012] Optionally, the set of parameters to be optimized for the deep autoregressive model includes the learning rate, the number of network layers, and the number of neurons.

[0013] Optionally, the step of generating a probability prediction interval for dam deformation using the optimized deep autoregressive model, and formulating dam deformation safety monitoring indicators based on the probability prediction interval, includes: The quantile prediction value corresponding to the upper quantile of the optimized deep autoregressive model is used as the upper limit of the probability prediction interval, and the quantile prediction value corresponding to the lower quantile is used as the lower limit of the probability prediction interval, thus obtaining the probability prediction interval under the preset confidence level. Calculate the coverage probability, average width, and coverage width criterion of the probability prediction interval, and evaluate the quality of the probability prediction interval; The upper limit of the probability prediction interval is used as the dam deformation safety monitoring index under the preset confidence level.

[0014] Optionally, it also includes: The distribution fit test is performed on the prediction error sequence of the optimized deep autoregressive model to determine whether the prediction error sequence follows a normal distribution. If the prediction error sequence follows a normal distribution, then the confidence interval at the same confidence level is calculated according to the 3σ principle, and the upper limit of the confidence interval is used as a comparison monitoring indicator. The dam deformation safety monitoring index, which is based on the probability prediction interval, is compared and analyzed with the comparative monitoring index, which is based on the confidence interval method, to evaluate the performance of the two monitoring indices in terms of conservatism and differences.

[0015] Optionally, the step of performing a distribution fit test on the optimized deep autoregressive model's prediction error sequence to determine whether the prediction error sequence follows a normal distribution includes: Arrange the sample data of the prediction error sequence in ascending order to construct an empirical distribution function; Calculate the maximum absolute difference between the empirical distribution function and the theoretical normal distribution function; The maximum absolute difference is compared with a preset threshold value; If the maximum absolute difference is less than the preset critical value, then the prediction error sequence is determined to follow a normal distribution.

[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method for formulating dam deformation safety monitoring indicators based on factor screening. By acquiring dam deformation monitoring data and influencing factors and constructing an initial factor set, redundant interference factors can be eliminated, reducing data noise and input dimensionality. By screening the correlation of influencing factors, effective factors closely related to deformation can be retained, improving the quality of model input. By using kernel linear discriminant analysis to extract principal component factors, the data dimensionality can be further compressed while retaining key influencing information, improving model computational efficiency and adapting to nonlinear correlation characteristics. By constructing an improved deep autoregressive model based on quantile regression, probabilistic prediction of dam deformation can be achieved, better fitting non-stationary and non-Gaussian monitoring data. By optimizing model parameters using the walrus optimization algorithm, automatic parameter optimization can be achieved to improve prediction accuracy and stability. By using the optimized probabilistic prediction model to generate probabilistic prediction intervals and formulate monitoring indicators, the shortcomings of traditional methods, such as insufficient samples, poor indicator reliability, and difficulty in reflecting deformation uncertainty, can be overcome. Ultimately, the method achieves accurate, reliable, and scientific formulation of dam deformation safety monitoring indicators, significantly improving the engineering applicability and safety assurance capability of the indicators. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating the method for determining dam deformation safety monitoring indicators based on factor screening, as provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the impact factor correlation screening and KLDA dimensionality reduction process provided in the embodiments of this application. Figure 3 A schematic diagram of the structure of the DeepAR model with improved quantile regression provided in the embodiments of this application; Figure 4 This is a schematic diagram illustrating the process of optimizing model parameters using the walrus optimization algorithm provided in this embodiment of the application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] This application provides a method for formulating dam deformation safety monitoring indicators based on factor screening. This method is executed by computer equipment, specifically by a terminal or server alone, or by both a terminal and a server.

[0022] Please see Figure 1 In an exemplary embodiment, the method for formulating dam deformation safety monitoring indicators based on factor screening includes: S110. Obtain dam deformation monitoring data and its corresponding influencing factors, and construct an initial factor set.

[0023] For example, dam deformation monitoring data can be deformation data from monitoring points. An initial factor set can be constructed by analyzing the deformation and influencing factor monitoring data of the corresponding monitoring points. This initial factor set can include three traditional deformation factor models: HTT, HST, and HT. A T.

[0024] S120. Perform correlation analysis on each influencing factor in the initial factor set, and select the influencing factors that meet the correlation threshold according to the preset correlation threshold.

[0025] For example, step S120 above (performing correlation analysis on each influencing factor in the initial factor set and selecting influencing factors that meet the correlation threshold according to a preset correlation threshold) may include: S121. The Spearman rank correlation coefficient method is used to calculate the rank correlation coefficient between each influencing factor in the initial factor set and the dam deformation monitoring data.

[0026] Specifically, using the HTT, HST, and HT constructed in step S110 A The traditional deformation factor model is used to perform correlation analysis between the various influencing factors in each factor model and the dam deformation monitoring data based on the Spearman rank correlation coefficient method, and the rank correlation coefficient is calculated.

[0027] S122. Based on the magnitude of the rank correlation coefficient, each influencing factor is divided into several different correlation levels.

[0028] Specifically, based on the calculated rank correlation coefficient, each influencing factor is divided into five correlation levels: extremely strong correlation (rank correlation coefficient 0.8~1.0), strong correlation (0.6~0.8), moderate correlation (0.4~0.6), weak correlation (0.2~0.4), and extremely weak correlation or no correlation (0~0.2).

[0029] S123. Based on the preset correlation threshold, remove the influence factors corresponding to the correlation levels that do not reach the correlation threshold, and obtain the influence factors that meet the correlation threshold.

[0030] Specifically, using a rank correlation coefficient of 0.2 as the preset correlation threshold, factors with extremely weak correlation or no correlation (rank correlation coefficient < 0.2) are screened out, while the remaining factors are retained as the factor set after correlation screening.

[0031] The formula for calculating the Spearman rank correlation coefficient is as follows: (1) In the formula: It is rank-correlated sparsity. It is order, It is a rank sequence. It refers to the number of samples.

[0032] S130. Kernel linear discriminant analysis was used to reduce the dimensionality of the screened influencing factors and extract the principal component factors.

[0033] For example, step S130 above (using kernel linear discriminant analysis to reduce the dimensionality of the screened influence factors and extract principal component factors) may include: S131. Using the selected impact factors as input, the selected impact factors are mapped to a high-dimensional feature space through the Gaussian kernel function.

[0034] S132. In a high-dimensional feature space, the optimal projection direction is found by maximizing the ratio of inter-class divergence to intra-class cohesion. The optimal projection direction is the projection direction that maximizes the inter-class separation and intra-class cohesion of the mapped samples in the low-dimensional space.

[0035] Among them, the inter-class scatter matrix is ​​used to measure the degree of dispersion between samples of different classes, and the intra-class clustering matrix is ​​used to measure the degree of clustering within samples of the same class. Kernel linear discriminant analysis (KLDA) is used to project high-dimensional space samples to the optimal direction to achieve effective dimensionality reduction.

[0036] S133. Project the selected influence factors onto the low-dimensional space according to the optimal projection direction to obtain the dimensionality-reduced influence factors. For multi-class linear discrimination scenarios, the projection into the low-dimensional space is represented by a hyperplane, and the projection weights can be represented by coefficient vectors. Solving for the projection weight vectors is transformed into solving for the eigenvectors of the corresponding non-zero eigenvalues ​​of the matrix.

[0037] S134. Plot the cumulative contribution rate curve, and extract the principal component factors from the dimensionality-reduced influence factors whose cumulative contribution rate reaches the preset cumulative contribution rate threshold, ensuring that the principal component factors can completely retain the core information affecting dam deformation.

[0038] For example, the cumulative contribution rate threshold can be 95%.

[0039] Specifically, kernel linear discriminant analysis (KLDA) uses linear discriminant analysis to find the optimal projection direction, maximizing the ratio of between-class divergence to intra-class cohesion of the mapped samples. The calculation formula is as follows: (2) (3) (4) In the formula: For higher-dimensional space In low-dimensional space The matrix formed by the projected vectors; The inter-class scatter matrix; This is the class cohesion matrix; For the first Class samples in high-dimensional space The dataset; For the first Class samples in high-dimensional space The mean vector; For higher-dimensional space The mean vector of all samples.

[0040] For multi-class linear discrimination, the low-dimensional space is projected onto a hyperplane, which transforms the matrix... It can be written in the following dot product form: (5) In the formula: This is a coefficient vector used to represent the projection weights.

[0041] Substituting the inter-class scatter matrix, intra-class clustering matrix, and dot product form into the projection objective function yields the projection weight vector that needs to be optimized in KLDA. The solution formula is for The solution is transformed into finding the eigenvectors corresponding to the non-zero eigenvalues ​​of the matrix; the projection of any sample in space onto the projection direction is calculated as follows: For example, substituting equations (3) to (5) into equation (1), we get: (6) In the formula: This refers to the projection weight vector that needs to be optimized in KLDA. and for 1-order matrix.

[0042] right The solution is transformed into the solution The eigenvectors corresponding to the non-zero eigenvalues. For any sample in the space. ,exist space The projection of the direction is calculated as follows: (7) For example, a Gaussian kernel function is used to implement the kernel transformation. The expression for the Gaussian kernel function is: (8) In the formula: For kernel parameters.

[0043] The cumulative contribution rate threshold is preferably around 95%. Principal components with a cumulative contribution rate reaching this threshold are selected as input variables for subsequent prediction models to fully preserve the core impact information of dam deformation.

[0044] S140. Based on the extracted principal component factors, construct a deep autoregressive model based on quantile regression improvement, and optimize the parameters of the deep autoregressive model.

[0045] Specifically, step S140 (constructing a deep autoregressive model based on quantile regression improvement based on the extracted principal component factors, and optimizing the parameters of the deep autoregressive model) may include: S141. Using the extracted principal component factors as input variables, a deep autoregressive model based on a recurrent neural network is constructed. The deep autoregressive model is used to describe the probability distribution of dam deformation monitoring data over time.

[0046] S142. Replace the output of the deep autoregressive model with the distribution parameters and the predicted quantile values ​​corresponding to multiple preset quantiles. The preset quantiles include upper quantiles and lower quantiles.

[0047] S143. The quantile loss function is used as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function. The model parameters are learned by minimizing the quantile loss function.

[0048] Specifically, For a given dam deformation past time series and covariates Deep autoregressive models are used to build future time series. The conditional probability distribution is as follows: (9) In the formula: Time series In time The value at; These are the points in time that are divided.

[0049] The above probability distribution can be transformed into likelihood form: (10) (11) In the formula: It is the likelihood function; For likelihood function Parameters; These are network structure parameters; Output for hidden layer; It is a multi-layer recurrent network containing RNNs.

[0050] Model training is achieved by maximizing the log-likelihood probability. Specifically, during training, DeepAR takes the previous hidden layer as input. Target value and covariates at the current time and model parameters This will allow you to obtain the hidden layer output at the current time. Thus, the likelihood function is calculated. parameters Finally, the optimal model parameters are learned by maximizing the log-likelihood probability. .

[0051] (12) In the formula: It is the log-likelihood function; Sampling time.

[0052] Furthermore, the distribution parameter output of the deep autoregressive model is replaced with multiple quantile predictions. ,in The target quantile is used; the quantile loss function is used instead of the negative log-likelihood function, and its expression is: (13) In the formula: These are the measured values ​​of dam deformation. These are the predicted quantile values; This represents the number of samples.

[0053] Furthermore, step S141 above (using the quantile loss function as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function, and learning the model parameters by minimizing the quantile loss function) may include: During the training process, the input of the deep autoregressive model at each time step includes the hidden layer output of the previous time step, the dam deformation monitoring data of the previous time step, and the covariates of the current time step. The hidden layer output of the current time step is obtained through the calculation of the recurrent neural network, and then the quantile loss function value of the current time step is calculated based on the hidden layer output of the current time step. By minimizing the sum of the quantile loss function values ​​of all time steps, the optimal parameters of the deep autoregressive model are learned.

[0054] Furthermore, the parameters of the deep autoregressive model are optimized using the walrus optimization algorithm, including: S1401. Initialize the walrus population and parameters.

[0055] Initialize the population and determine the position of each walrus individual in the population. The position of each walrus individual represents a set of parameters to be optimized in the depth autoregressive model.

[0056] S1402, Location update during foraging phase.

[0057] During the foraging phase, each walrus individual updates its own position under the guidance of the currently best individual in order to search for a better combination of parameters.

[0058] S1403, Location update during migration phase.

[0059] During the migration phase, each walrus individual randomly selects another individual in the population and decides to move closer to or further away from that individual based on the selected individual's fitness, in order to expand the search area.

[0060] S1404, Local search during the predator evasion phase.

[0061] During the predator evasion phase, each walrus individual conducts a localized search within a small area near its current location to fine-tune its parameter combinations.

[0062] S1405. Determine if the iteration termination condition has been met; if not, continue iterating; if so, output the optimal parameter combination.

[0063] The process iteratively executes the foraging phase, the migration phase, and the predator evasion phase until the termination condition is met, and outputs the location of the optimal walrus individual. The location of the optimal walrus individual is the optimal parameter combination of the depth autoregressive model.

[0064] Specifically, the walrus optimization algorithm initializes the population by generating individuals using a uniform distribution, as shown in the formula: (14) In the formula: Let this be the position of the i-th walrus individual; A random number between 0 and 1; , These are the upper and lower bounds of the population space, respectively.

[0065] Foraging phase location update formula: (15) (16) In the formula: For the first The original location of the walrus individual dimension; For the first The original fitness value of an individual walrus; The first stage of walrus foraging The location where each walrus individual is newly formed For its first dimension, Its fitness value; A random number between 0 and 1; The current optimal individual, i.e., the leader. Its dimensional; A random number between 1 and 2.

[0066] Migration phase position update formula: (17) (18) In the formula: The first stage of walrus migration The newly generated position of each individual For its first dimension, Its fitness value; Another walrus individual was randomly selected. Its fitness value.

[0067] Position update formula for the predator phase: (19) (20) (twenty one) In the formula: The first stage for walruses to escape or fight predators The newly generated position of each individual For its first dimension, Its fitness value; This represents the current iteration number; , For the individual at the current iteration number... Upper and lower bounds for changes in position in each dimension; , The upper and lower bounds of the population space are respectively the first. dimension.

[0068] The iteration terminates when the maximum number of iterations or the fitness convergence condition is met, and the optimal parameter combination is output.

[0069] Optionally, a set of parameters to be optimized for a deep autoregressive model may include key parameters that affect the model’s prediction accuracy and convergence speed, such as the learning rate, the number of network layers, and the number of neurons.

[0070] S150. The optimized deep autoregressive model is used to generate the probability prediction interval of dam deformation, and the dam deformation safety monitoring index is formulated based on the probability prediction interval.

[0071] In an exemplary embodiment, step S150 (generating a probability prediction interval for dam deformation using the optimized deep autoregressive model, and formulating dam deformation safety monitoring indicators based on the probability prediction interval) includes: First, the predicted quantile values ​​of the upper quantiles output by the optimized deep autoregressive model are used as the upper limit of the probability prediction interval, and the predicted quantile values ​​of the lower quantiles are used as the lower limit of the probability prediction interval, thus obtaining the probability prediction interval under the preset confidence level. 2. Calculate the coverage probability, average width, and coverage width criterion of the probability prediction interval, and evaluate the quality of the probability prediction interval; Third, the upper limit of the probability prediction interval is used as the dam deformation safety monitoring indicator under the pre-set confidence level.

[0072] Specifically, in step one, the preset confidence level is preferably 99%, with the corresponding upper quantile being 0.995 and the lower quantile being 0.005. The 0.995 quantile prediction value is used as the upper limit of the probability prediction interval, and the 0.005 quantile prediction value is used as the lower limit of the probability prediction interval, thus obtaining the dam deformation probability prediction interval at a 99% confidence level.

[0073] Step two uses the predicted interval coverage probability (I) PICP ), average width of the prediction interval (I) PINAW ), Comprehensive width coverage criterion (I) WC The three indicators are used to quantitatively evaluate the probability prediction interval, and the calculation formula is as follows: (twenty two) (twenty three) (twenty four) In the formula: This represents the number of samples for predicting dam deformation. The target value range; It is a Boolean value; , These represent the upper and lower limits of the deformation prediction range, respectively.

[0074] In an exemplary embodiment, the method for determining dam deformation safety monitoring indicators based on factor screening further includes: The distribution fit test is performed on the prediction error sequence of the optimized deep autoregressive model to determine whether the prediction error sequence follows a normal distribution. If the prediction error sequence follows a normal distribution, then the confidence interval at the same confidence level is calculated according to the 3σ principle, and the upper limit of the confidence interval is used as a comparison and monitoring indicator. The dam deformation safety monitoring index proposed based on probability prediction intervals is compared and analyzed with the comparative monitoring index proposed based on the confidence interval method to evaluate the performance of the two monitoring indices in terms of conservatism and differences.

[0075] Furthermore, a distribution fit test is performed on the prediction error sequence of the optimized deep autoregressive model to determine whether the prediction error sequence follows a normal distribution, including: Arrange the sample data of the prediction error sequence in ascending order to construct an empirical distribution function; Calculate the maximum absolute difference between the empirical distribution function and the theoretical normal distribution function; The maximum absolute difference is compared with a preset threshold value; If the maximum absolute difference is less than the preset critical value, then the prediction error sequence is judged to follow a normal distribution.

[0076] Specifically, the probability density function test method is used to arrange the sample data of the prediction error sequence in ascending order and remove duplicate values, thus dividing the sample interval into... Divide the sample data into equidistant and non-overlapping intervals, count the frequency of the sample data falling into each interval, and analyze the error distribution pattern. Construction of the empirical distribution function: Let the prediction error sequence be... The sample size is Sort the error samples from smallest to largest as follows: For any real number Constructing the empirical distribution function Then we have: (25) Using the Kolmokolov-Smilov test, the null hypothesis H0 is proposed: the error sequence follows a pre-defined theoretical distribution, that is, the empirical distribution function of the sample is consistent with the theoretical distribution function (i.e., ).

[0077] Calculate the absolute difference between the empirical distribution function value Fn(x) and the theoretical distribution function value F(x), and take the maximum value as the test statistic Dn; where, .

[0078] The significance level is preferably 0.01, depending on the sample size. Significance level The critical value is obtained by looking up the table. ; the maximum absolute difference With preset threshold Compare; if If so, then the null hypothesis is accepted, and the prediction error sequence is judged to follow a normal distribution.

[0079] If the prediction error sequence follows a normal distribution, a confidence interval is determined based on the 3σ principle (for example, the 2.576σ principle) and a 99% confidence level. The upper limit of this confidence interval is then used as the comparative monitoring indicator for the confidence interval method, thus completing the comparison and applicability evaluation of the monitoring indicators of the two methods.

[0080] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0081] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0082] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0084] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for formulating dam deformation safety monitoring indicators based on factor screening, characterized in that, include: Acquire dam deformation monitoring data and their corresponding influencing factors, and construct an initial factor set; Correlation analysis is performed on each influencing factor in the initial factor set, and influencing factors that meet the preset correlation threshold are selected. Kernel linear discriminant analysis was used to reduce the dimensionality of the screened influencing factors and extract principal component factors. Based on the extracted principal component factors, a deep autoregressive model based on quantile regression is constructed, and the parameters of the deep autoregressive model are optimized. The optimized deep autoregressive model is used to generate a probability prediction interval for dam deformation, and safety monitoring indicators for dam deformation are formulated based on the probability prediction interval.

2. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 1, characterized in that, The step of performing correlation analysis on each influencing factor in the initial factor set, and filtering out influencing factors that meet the preset correlation threshold according to a preset correlation threshold, includes: The Spearman rank correlation coefficient method was used to calculate the rank correlation coefficient between each of the influencing factors in the initial factor set and the dam deformation monitoring data. Based on the magnitude of the rank correlation coefficient, each of the influencing factors is divided into several different correlation levels; Based on the preset relevance threshold, influence factors corresponding to the relevance levels that do not reach the relevance threshold are removed, and influence factors that meet the relevance threshold are obtained.

3. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 1, characterized in that, The screened influencing factors were subjected to dimensionality reduction using kernel linear discriminant analysis to extract principal component factors, including: The selected impact factors are used as input, and the Gaussian kernel function is used to map the selected impact factors to a high-dimensional feature space. In the high-dimensional feature space, the optimal projection direction is found by maximizing the ratio of inter-class divergence to intra-class cohesion. The optimal projection direction is the projection direction that maximizes the inter-class separation and intra-class cohesion of the mapped samples in the low-dimensional space. The selected influence factors are projected into a low-dimensional space according to the optimal projection direction to obtain the dimension-reduced influence factors. Plot the cumulative contribution rate curve, and extract the principal component factors from the dimensionality-reduced influence factors whose cumulative contribution rate reaches the preset cumulative contribution rate threshold.

4. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 1, characterized in that, The process involves constructing a deep autoregressive model based on quantile regression improvements using the extracted principal component factors, and optimizing the parameters of the deep autoregressive model, including: Using the extracted principal component factors as input variables, a deep autoregressive model based on a recurrent neural network is constructed. The deep autoregressive model is used to describe the probability distribution of dam deformation monitoring data over time. The output of the deep autoregressive model is replaced by the distribution parameters with the quantile prediction values ​​corresponding to multiple preset quantiles, wherein the preset quantiles include upper quantiles and lower quantiles; The quantile loss function is used as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function. The model parameters are learned by minimizing the quantile loss function.

5. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 4, characterized in that, The method employs a quantile loss function as the optimization objective for training the deep autoregressive model, replacing the original negative log-likelihood function. The model parameters are learned by minimizing the quantile loss function, including: learning model parameters using the loss function. During the training process, the input of the deep autoregressive model at each time step includes the hidden layer output of the previous time step, the dam deformation monitoring data of the previous time step, and the covariates of the current time step. The hidden layer output of the current time step is obtained through the calculation of the recurrent neural network, and then the quantile loss function value of the current time step is calculated based on the hidden layer output of the current time step. By minimizing the sum of the quantile loss function values ​​of all time steps, the optimal parameters of the deep autoregressive model are learned.

6. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 4, characterized in that, The parameter optimization of the deep autoregressive model, using the walrus optimization algorithm, includes: Initialize the population and determine the position of each walrus individual in the population. The position of each walrus individual represents a set of parameters to be optimized in the depth autoregressive model. During the foraging phase, each walrus individual updates its own position under the guidance of the current best individual in order to search for a better combination of parameters; During the migration phase, each walrus individual randomly selects another individual in the population and decides to move closer to or further away from that individual based on the fitness of the selected individual in order to expand the search range; During the predator evasion phase, each walrus individual conducts a local search within a small area around its current location to fine-tune its parameter combinations; The foraging phase, migration phase, and predator evasion phase are executed iteratively until the termination condition is met, and the location of the optimal walrus individual is output. The location of the optimal walrus individual is the optimal parameter combination of the depth autoregressive model.

7. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 6, characterized in that, The set of parameters to be optimized for the deep autoregressive model includes the learning rate, the number of network layers, and the number of neurons.

8. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 1, characterized in that, The process involves generating a probability prediction interval for dam deformation using the optimized deep autoregressive model, and formulating dam deformation safety monitoring indicators based on the probability prediction interval, including: The quantile prediction value corresponding to the upper quantile of the optimized deep autoregressive model is used as the upper limit of the probability prediction interval, and the quantile prediction value corresponding to the lower quantile is used as the lower limit of the probability prediction interval, thus obtaining the probability prediction interval under the preset confidence level. Calculate the coverage probability, average width, and coverage width criterion of the probability prediction interval, and evaluate the quality of the probability prediction interval; The upper limit of the probability prediction interval is used as the dam deformation safety monitoring index under the preset confidence level.

9. The method for formulating dam deformation safety monitoring indicators based on factor screening according to claim 1, characterized in that, Also includes: The distribution fit test is performed on the prediction error sequence of the optimized deep autoregressive model to determine whether the prediction error sequence follows a normal distribution. If the prediction error sequence follows a normal distribution, then the confidence interval at the same confidence level is calculated according to the 3σ principle, and the upper limit of the confidence interval is used as a comparison monitoring indicator. The dam deformation safety monitoring index, which is based on the probability prediction interval, is compared and analyzed with the comparative monitoring index, which is based on the confidence interval method, to evaluate the performance of the two monitoring indices in terms of conservatism and differences.

10. The method for determining dam deformation safety monitoring indicators based on factor screening according to claim 9, characterized in that, The step of performing a distribution fit test on the prediction error sequence of the optimized deep autoregressive model to determine whether the prediction error sequence follows a normal distribution includes: Arrange the sample data of the prediction error sequence in ascending order to construct an empirical distribution function; Calculate the maximum absolute difference between the empirical distribution function and the theoretical normal distribution function; The maximum absolute difference is compared with a preset threshold value; If the maximum absolute difference is less than the preset critical value, then the prediction error sequence is determined to follow a normal distribution.