Underground powerhouse seepage characteristic factor set construction method considering hysteresis effect
By screening hysteresis characteristic factors using a Bayesian vector autoregression model and impulse response function, and combining the ReliefF algorithm and variance expansion coefficient, a high-precision seepage prediction model is constructed. This solves the problem of insufficient prediction accuracy caused by ignoring hysteresis effects in traditional models, and achieves more accurate dynamic response analysis of the seepage field.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional seepage statistical models cannot effectively capture the lag effect caused by the heterogeneity of spatial location and geological conditions during seepage in underground power plants, resulting in insufficient model universality and prediction accuracy. Inappropriate selection of feature factors can lead to overfitting or poor fitting results.
By employing a Bayesian vector autoregression model combined with impulse response function analysis, hysteresis characteristic factors were screened out. Collinearity tests were conducted using the ReliefF algorithm and variance expansion coefficient to construct a high-precision seepage prediction model characteristic factor set.
It significantly improves the universality and prediction accuracy of the seepage prediction model, enhances the interpretability and generalization ability of the feature factors, and solves the problem of ignoring the influence of hysteresis effect in seepage prediction.
Smart Images

Figure CN121834754A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of seepage behavior monitoring and analysis of pumped storage power station underground powerhouse, and particularly relates to a method for constructing a set of seepage characteristic factors of underground powerhouse considering hysteresis effect. BACKGROUND
[0002] The seepage problem of pumped storage power station underground powerhouse is a complex physical process under the coupling of multiple factors, and there is a significant time lag effect in its dynamic response. Restricted by the time and space propagation characteristics of the seepage process, whether it is the migration of water, the transmission of pore water pressure, or the change of temperature field, it needs to pass through a certain time to be transmitted from the boundary to the powerhouse area in the highly heterogeneous and anisotropic fractured rock mass medium. This hysteresis effect makes the influence of environmental factors on the seepage problem not instantaneous, but presents a significant lagging characteristic, and the hysteresis effect research becomes a key and difficult point in the construction of seepage prediction model and safety monitoring analysis of pumped storage power station underground powerhouse.
[0003] In the traditional seepage statistical model, the fixed early average component is usually used to describe the hysteresis effect existing in the seepage process. For water level component and rainfall component, the average values of the first, second, fifth and tenth days before the monitoring day are usually selected as the characteristic factors; for temperature component, a sine or cosine periodic function is usually directly used for fitting. Although this method describes the hysteresis effect existing in the seepage process to some extent, in actual engineering application, the seepage problem is significantly dependent on various factors such as hydrogeological conditions of specific project, development and connectivity degree of rock mass fracture network, and spatial position of water source relative to monitoring point. The response of a monitoring point near a strong permeable fracture zone may be extremely rapid (short lag days); while a deep monitoring point surrounded by dense rock mass or complex fracture network may show a significant lag of several weeks or even months. The traditional method cannot capture this dynamic response law varying with spatial position and geological conditions, ignores the external specificity and internal delay of the seepage system, and thus leads to poor universality and prediction accuracy of the model when applied across projects or monitoring points, making it difficult to accurately depict the dynamic response of the real seepage field.
[0004] Meanwhile, the number of characteristic factors in the set of characteristic factors also has a great influence on the fitting effect of the prediction model. If too many characteristic factors are introduced, although the fitting goodness of the training sample can be improved, irrelevant or redundant features may be introduced, leading to feature dimension expansion and multicollinearity problem, causing model overfitting, and reducing prediction accuracy and generalization ability. On the contrary, if there are too few characteristic factors, the key driving factors of the seepage system cannot be fully captured, which also leads to poor fitting effect and prediction performance of the model. Therefore, the selection of characteristic factors needs to seek an optimal balance between fitting goodness and model complexity. SUMMARY
[0005] The present application aims at the problems in the background art, and provides a set construction method of seepage characteristic factors of underground powerhouse considering hysteresis effect, which quantitatively analyzes the hysteresis effect of influencing factors such as water pressure and temperature based on long-term monitoring data of actual engineering seepage, obtains dynamic hysteresis factors accurately reflecting spatial and temporal variability, and adopts a characteristic screening method to eliminate redundant and secondary factors to construct a set of seepage prediction model characteristic factors of pumped storage power station underground powerhouse considering hysteresis effect, thereby providing reliable characteristic input for constructing a high-precision seepage prediction model of pumped storage power station underground powerhouse.
[0006] To achieve the above-mentioned purpose, the technical scheme of the present application is as follows: a set construction method of seepage characteristic factors of underground powerhouse considering hysteresis effect, comprising:
[0007] Step S1, set construction of seepage prediction model characteristic factors of pumped storage power station underground powerhouse:
[0008] By analyzing the seepage mechanism and influencing factors of pumped storage power station underground powerhouse, an initial set of characteristic factors covering reservoir pressure, temperature, rainfall and time effect factors is obtained;
[0009] Step S2, parameter solving of hysteresis characteristic factors:
[0010] By establishing a Bayesian vector autoregressive model, the hysteresis days and influence days of factors including reservoir outlet water level and temperature on seepage are solved by applying impulse response function analysis, and hysteresis characteristic factors are obtained;
[0011] Step S3, importance screening of characteristic factors:
[0012] By performing factor contribution degree analysis through the ReliefF algorithm, performing collinearity test by using the variance inflation coefficient, combining factor importance ranking and comprehensive screening, an optimized set of characteristic factors with high prediction performance, strong interpretability and good generalization ability is finally determined, thereby providing reliable characteristic input for constructing a high-precision seepage prediction model of underground powerhouse.
[0013] Further, in step S1, the initial set of characteristic factors is represented as:
[0014]
[0015] Among them, , , , respectively represent the water pressure component, the temperature component, the rainfall component and the time effect component.
[0016] Further, step S1 is specifically:
[0017] Step S11, monitoring data preprocessing:
[0018] The original data of the seepage of the underground powerhouse of the pumped storage power station is preprocessed, and the missing values and abnormal values in the original data are repaired; meanwhile, the data of different sources are standardized in format, and the multi-source heterogeneous data are combined into a unified format.
[0019] Step S12, seepage characteristic factor set construction:
[0020] The influence factors of the seepage effect of the underground powerhouse of the pumped storage power station are analyzed, and an initial characteristic factor set of the seepage prediction model of the underground powerhouse of the pumped storage power station is constructed.
[0021] Further, in step S11, for the missing value processing, a cubic spline interpolation method is adopted.
[0022] Further, in step S11, for the selection of the monitoring data sampling area, the distribution of the seepage monitoring points of the underground powerhouse of the pumped storage power station, the engineering geological conditions and the potential seepage path are selected, and the original monitoring points and the actual needs are combined to determine comprehensively.
[0023] Further, in step S11, the sampling time should cover at least one complete hydrological year, and include the characteristic working conditions of the reservoir operation period: flood period, discharge period and dry period.
[0024] Further, step S2 is specifically:
[0025] Step S21, establishing a BVAR Bayesian vector autoregressive model:
[0026] The BVAR model can analyze the dynamic law that the dependent variable is affected by the independent variables at past time, effectively regularize the model parameters with prior information, reduce the number of parameters to be estimated without introducing subjective factors. It has obvious advantages in capturing the correlation and long-term effect of time series data. The BVAR Bayesian vector autoregressive model is expressed as:
[0027]
[0028] In the formula: is a state vector of dimension , representing the value of different variables at time , is the number of variables; is the lag order of ; is a constant vector representing the stable value of the state without any external influence; is the number of samples; is a dimension coefficient matrix describing the influence of the past vector on the current vector ; The disturbance vector represents the difference between the actual value and the predicted value caused by random disturbance or model error.
[0029] Step S22, ADF data stationarity test:
[0030] The BVAR model requires that the time series be stationary, and data stationarity test is needed. ADF test is applied to determine whether the time series data has a unit root, to avoid the problem of "pseudo-regression", and to eliminate the trend and seasonality in the data. First-order difference formula is applied to convert the non-stationary time series into stationary time series. The core formula of ADF test is as follows:
[0031]
[0032] In the formula, is the monitoring value at time t; is the first-order difference, representing the daily change amount of the monitoring value, and is used to eliminate short-term fluctuations; is the constant term, reflecting the deviation of the sequence mean; is the time trend term; is the coefficient of the lag difference term, used to eliminate high-order autocorrelation; is the lag order, determined by information criterion; is the core parameter, used to determine whether the sequence is stationary; is the error term;
[0033] When the absolute value of the ADF statistic (the core parameter ) is less than the critical test value of the significance level, it indicates that the sequence is in a non-stationary state, and it needs to be processed by first-order difference to eliminate the trend and make the sequence stationary. The ADF test is performed again on the sequence after the difference, until the sequence becomes stationary.
[0034] Step S23, BIC information criterion to determine the optimal lag order:
[0035] Information criterion seeks a balance between model fitting degree and model simplicity. Its goal is to select the model with the smallest value, i.e., the maximum posterior probability, as the optimal model:
[0036]
[0037] In the formula, is the maximum likelihood estimate of the model, indicating the fitting degree of the model to the data; is the number of parameters (degrees of freedom) of the model, reflecting the complexity of the model; is the sample size, which increases with the number of parameters and sample size The increase, The penalties are also increasing;
[0038] Step S24: Calculate the hysteresis parameters using the impulse response function (IRF):
[0039] The impulse response function (IRF) quantifies the intensity and direction of the impact of a unit shock on other variables in a Bayesian vector autoregressive (BVAR) model at different lag times. It describes the series of effects on other variables (such as seepage) over a current and future period when a variable (e.g., reservoir water level) is subjected to a unit-sized "shock" or change. The IRF is expressed as:
[0040]
[0041] In the formula, The response value of the affected variable; For the affected variables; For the source of the impact; for exist The unit impact applied at any given time; For response time;
[0042] The lag days and impact days obtained from the impulse response function (IRF) analysis are used as key parameters and substituted into the lag effect function to obtain the lag characteristic factor.
[0043] Furthermore, step S3 specifically includes:
[0044] Step S31, Factor Contribution Analysis:
[0045] Too many influencing factors can lead to feature redundancy and cause overfitting in prediction accuracy. By selecting factors, the number of factors can be reduced, thereby improving the interpretability and generalization ability of the model. The algorithm introduces the concept of weight calculation, and achieves feature selection by considering the weight differences between neighboring samples of different categories. The specific steps of the algorithm are as follows:
[0046] (1) Initialize the number of iterations and the number of nearest neighbors selected ;
[0047] (2) Randomly select objects from the dataset From and Select from similar objects The nearest neighbor is denoted as ,and Selecting from heterogeneous objects The nearest neighbor is denoted as , repeat until the iteration sampling number is ;
[0048] (3) Calculate the distance between two objects on the characteristics :
[0049]
[0050] wherein, and are the values of the objects and on the characteristics ; and are the maximum and minimum values of the characteristics ;
[0051] The importance value :
[0052]
[0053] wherein, is the iteration number, ; is the probability of the object belonging to the same category ; is the probability of the object belonging to a different category ;
[0054] (4) Repeat steps (1)-(3) to output the importance vector , then the feature weight corresponding to the characteristics is:
[0055]
[0056] Step S32, VIF collinearity test:
[0057] When there is serious multicollinearity between factors, the independent variables weaken the ability to explain the contribution and change each other, thereby affecting the accuracy of the prediction result. The variance inflation coefficient VIF is used to test the multicollinearity of the factors, and the calculation formula is as follows:
[0058]
[0059] In the formula: represents the fitting coefficient of the independent variable;
[0060] The multicollinearity of the variable is judged by the size of the VIF value, the larger the VIF value, the stronger the multicollinearity between factors;
[0061] Step S33, characteristic factor screening:
[0062] According to the results of steps S31 and S32, in combination with the importance ranking and comprehensive screening, the factors with low importance in the high VIF factors are removed, and finally the optimized characteristic factor set with high prediction performance, strong interpretability and good generalization ability is determined, thereby providing reliable characteristic input for constructing a high-precision seepage prediction model of an underground powerhouse.
[0063] The application further provides a system for constructing a seepage characteristic factor set of an underground powerhouse considering hysteresis effect, which comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor, and when the processor executes the computer program instructions, the method described above can be realized.
[0064] The application further provides a computer readable storage medium, which stores computer program instructions capable of being executed by a processor, and when the processor executes the computer program instructions, the method described above can be realized.
[0065] Compared with the prior art, the application has the following beneficial effects: the method effectively solves the problem that the traditional method uses fixed early average components to describe the hysteresis effect of seepage, and ignores the differences in hydrogeological structure, development and connectivity of rock mass fracture network, spatial position of monitoring points relative to water sources and other factors, thereby leading to poor universality of model characteristic factors, low prediction accuracy and difficulty in accurately reflecting the influence mechanism of the real seepage field; the method significantly enhances the universality and interpretability of the characteristic factor set of the seepage prediction model, and improves the prediction accuracy of the seepage prediction model. BRIEF DESCRIPTION OF DRAWINGS
[0066] Figure 1 A flowchart of a method for constructing a seepage prediction model characteristic factor set of a pumped storage power station underground powerhouse considering hysteresis effect according to the application;
[0067] Figure 2 An environmental quantity monitoring value in an embodiment of the method for constructing a seepage prediction model characteristic factor set of a pumped storage power station underground powerhouse considering hysteresis effect according to the application;
[0068] Figure 3 A water pressure component influence weight distribution function curve in the method for constructing a seepage prediction model characteristic factor set of a pumped storage power station underground powerhouse considering hysteresis effect according to the application;
[0069] Figure 4 A hysteresis parameter calculation response strength map in an embodiment of the method for constructing a seepage prediction model characteristic factor set of a pumped storage power station underground powerhouse considering hysteresis effect according to the application. DETAILED DESCRIPTION
[0070] The technical solutions of the present application will be described in detail below with reference to the drawings.
[0071] The present application provides a set construction method of seepage characteristic factors of underground powerhouse considering hysteresis effect, comprising:
[0072] Step S1, set construction of seepage prediction model characteristic factors of pumped storage power station underground powerhouse:
[0073] Through analyzing the seepage mechanism and influencing factors of pumped storage power station underground powerhouse, an initial characteristic factor set covering reservoir pressure, temperature, rainfall and time effect factors is obtained;
[0074] Step S2, hysteresis characteristic factor parameter solving:
[0075] Through establishing a Bayesian vector autoregressive model, the hysteresis days and influence days of factors including reservoir outlet water level and temperature on seepage are solved by applying impulse response function analysis, and the hysteresis characteristic factors are obtained.
[0076] Step S3, characteristic factor importance screening:
[0077] Through ReliefF algorithm for factor contribution degree analysis, collinearity test is carried out by using variance inflation coefficient, combined with factor importance sorting and comprehensive screening, finally the optimized characteristic factor set with high prediction performance, strong interpretability and good generalization ability is determined, which provides reliable characteristic input for constructing high-precision seepage prediction model of underground powerhouse.
[0078] The following is the specific implementation process of the present application.
[0079] Please refer to Figure 1 The present application is an example of a set construction method of seepage prediction model characteristic factors of pumped storage power station underground powerhouse considering hysteresis effect, comprising the following three steps:
[0080] Step S1, set construction of seepage prediction model characteristic factors of pumped storage power station underground powerhouse:
[0081] Step S11, monitoring data preprocessing:
[0082] The original monitoring data of seepage of pumped storage power station underground powerhouse are preprocessed, including reservoir water level, temperature, rainfall, monitoring point permeation pressure, etc., to repair missing values and abnormal values in the original data; at the same time, different source data are standardized in format, and multi-source heterogeneous data are combined into a unified format.
[0083] For missing value processing, the cubic spline interpolation method is adopted. Seepage process is a physical process, its change is usually continuous and relatively smooth, the core value of data is to reflect the trend and capture the abnormal. The purpose of interpolation is often to fill in the missing data, smooth the curve to better observe the trend, or provide the basis for subsequent analysis (such as derivative calculation of seepage gradient).
[0084] Cubic spline interpolation ensures the continuity of the second derivative, which means that the curve it produces is very smooth and has no sharp corners. This is most suitable for seepage, a physical process, because its rate of change (flow rate, water pressure gradient) is also continuously changing. Avoid the Runge phenomenon of high-order polynomial interpolation. Even if the monitoring points are unevenly distributed, reliable results can be produced. It can better maintain the original shape of the data and will not introduce additional oscillations that do not conform to the physical law.
[0085] For the selection of monitoring data sampling area, it is necessary to select according to the distribution of seepage monitoring points in the underground powerhouse of pumping and storage power station, engineering geological conditions and potential seepage path, and to determine comprehensively combining the original monitoring points and actual needs. The monitoring points should be arranged along the seepage path, leakage area, main caverns of pumping and storage power station underground powerhouse, and should be arranged in layers and zones to capture the spatial distribution characteristics of seepage field.
[0086] The sampling time should cover at least one complete hydrological year, and especially include the characteristic working conditions of reservoir operation period: such as flood period, discharge period, dry period. In the stage of rapid change of seepage pressure, the monitoring frequency should not be less than once a day to accurately capture the seepage response; In the stable operation period, the frequency can be appropriately reduced, but not less than once a week, to ensure the time sequence continuity and load coverage comprehensiveness of the data.
[0087] As shown in Figure 2 , combined with the actual monitoring of monitoring data, the embodiment determines that the time series data coverage period is from May 1, 2020 to December 1, 2023, with a total of 1310 continuous time observation nodes.
[0088] Step S12, seepage characteristic factor set construction:
[0089] The seepage state of the powerhouse surrounding rock of pumping and storage power station, as a water structure deeply buried in rock mass, is also affected by multiple environmental variables and time factors, and shows significant spatio-temporal correlation and lag characteristics in the mechanism of action. Referring to the construction idea of seepage observation statistical model of concrete dam and earth-rock dam, the seepage pressure in the surrounding rock and supporting system of pumping and storage power station can be regarded as being composed of water pressure component, rainfall component, temperature component and time effect component, that is:
[0090]
[0091] (1) Water pressure component
[0092] water pressure component It includes the direct impact of reservoir water level and the spatiotemporal lag effect, characterizing the direct effect of reservoir water pressure transmitted to the surrounding rock of the powerhouse through the rock mass fracture network, exhibiting a composite characteristic of nonlinearity, lag and spatial transmission.
[0093] Changes in reservoir water level directly affect the boundary head of the surrounding rock of the pumped-storage powerhouse. Due to the complex permeability and connectivity of the rock mass fracture network, the response of seepage pressure to changes in reservoir water level is often not a simple linear proportional relationship. Therefore, the reservoir water level is introduced... higher-order terms ( It captures nonlinear hydrostatic pressure response and characterizes the possible changes in hydraulic gradient and seepage path under high pressure.
[0094] The seepage field triggered by reservoir water level fluctuations is a dynamic physical process. The propagation of water pressure waves in fractured media takes time. Reservoir water level changes are slowly transmitted to the power plant area via the rock fracture network, influenced by the seepage path and the hysteresis effect controlled by rock permeability. This means that the current seepage pressure is correlated with the reservoir water level at a certain point in the past. Therefore, the reservoir water level is introduced... The lagged term ( This characterizes the lag effect of reservoir water level.
[0095] Traditional methods using the average water level in the early stages to replace the concept of the lag component of water pressure are rather vague. In reality, the influence of upstream water level on seepage may occur on the 2nd day, the 5th day, etc., and its effect does not disappear immediately thereafter. Its influence is a gradual process of rising and falling, not an average process. To reflect the effect of continuous water level changes on seepage, a weighted distribution function of the reservoir water level's influence is introduced. ,like Figure 3 As shown. According to statistical analysis, the influence of the weight distribution function... If the water level basically follows a normal distribution, then the lag component is:
[0096]
[0097] In the formula, The coefficients are constants. for The upstream water level at any given time; This refers to the number of days the upstream water level lagged behind. This represents the standard deviation of the normal distribution of upstream water levels, i.e., the number of days of impact.
[0098] The upstream water level is typically measured daily; therefore, the continuous integral can be converted to a discrete integral, and the integration interval only needs to be taken as... Two to three times the amount is sufficient to meet the requirements.
[0099] In addition, due to the continuity of seepage field in space, the measured value of the seepage pressure monitoring point located on the upstream side of the powerhouse and closer to the upstream reservoir reflects the water pressure state of the reservoir water level after a certain rock mass permeation path, providing intermediate state information from the source term to the powerhouse. Therefore, the measured value of the upstream monitoring point is introduced into the model. At the same time, is also the result of the time and space lagging effect of the reservoir water level. Therefore, the lag term of the measured value of the monitoring point is introduced, representing its leading influence on the powerhouse monitoring point.
[0100] In summary, the characteristic factor set of the water pressure component can be expressed as:
[0101]
[0102] (2) Temperature component
[0103] The seepage state of the underground powerhouse of the pumped storage power station is closely related to the temperature of the surrounding rock. When the temperature of the surrounding rock decreases, the rock mass produces thermal contraction, causing the original fissures to expand or new micro-fissures to form, thereby increasing the permeation channel and exacerbating the seepage phenomenon; on the contrary, when the temperature rises, the rock mass expands, causing some fissures to shrink and close, and the permeability accordingly decreases.
[0104] The temperature field of the underground powerhouse of the pumped storage power station changes slowly and is mainly driven by changes in the external environment temperature . Due to the low heat conduction speed of the surrounding rock, the seepage effect lags behind the changes in the external environment temperature, and short-term temperature fluctuations have little effect on seepage, while long-term temperature trends have a more significant impact on seepage effects. Therefore, the lag term of temperature is introduced, the average temperature value is used to smooth short-term fluctuations, and the long-term trend is highlighted, representing the lagging influence of the temperature field on seepage effects.
[0105]
[0106] In the formula, is a constant coefficient; are the average air temperatures on the monitoring day, the day before the monitoring day to the day before the monitoring day, the day before the monitoring day to the day before the monitoring day, and the day before the monitoring day to the day before the monitoring day, respectively; is the annual average temperature; , are the lag days and influence days of temperature on seepage.
[0107] If there is no measured bedrock temperature, a sinusoidal periodic function can be used as the temperature component.
[0108]
[0109] In summary, the temperature component characteristic factor set can be expressed as:
[0110]
[0111] (3) Rainfall component
[0112] Rainfall will increase the water content of rock-soil mass through surface infiltration, change the seepage field and pore water pressure distribution of surrounding rock. Since the migration of rainwater in the aeration zone and the infiltration in the rock mass fracture network need time, the seepage change in the powerhouse area has significant hysteresis compared to the rainfall process. The early rainfall method is used to consider its influence.
[0113]
[0114] In the formula, is a constant coefficient; is the average rainfall of the monitoring day, 1 day before the monitoring day, 2-4 days before the monitoring day, and 5-8 days before the monitoring day, respectively; is the average rainfall;
[0115] In summary, the rainfall component model can be expressed as:
[0116]
[0117] (4) Time component
[0118] The time component is an irreversible component that develops in an irreversible direction with the passage of time. Under the action of static water and gravity, the inherent complex creep of the surrounding rock of the underground powerhouse of the pumped storage power station may change the internal cracks and seepage paths of the surrounding rock. The influence of the time component usually shows a certain regularity, which changes sharply at the beginning and gradually stabilizes later. The time component can be expressed by a linear function and a logarithmic function:
[0119]
[0120] In the formula, is the number of days from the beginning of reservoir impoundment divided by 100.
[0121] In summary, the seepage model factor set of the underground powerhouse of the pumped storage power station is obtained:
[0122]
[0123] Step S2, solve the hysteresis characteristic factor parameters:
[0124] Step S21, establishing a BVAR Bayesian vector autoregressive model:
[0125] The BVAR model can analyze the dynamic law of the dependent variable affected by the independent variable at the past time, effectively regularize the model parameters with prior information, reduce the number of parameters to be estimated without introducing subjective factors. It has significant advantages in capturing the correlation and long-term effects of time series data. The BVAR model can be expressed as:
[0126]
[0127] In the formula: is a state vector of dimension , representing the values of different variables at time ; is the number of variables; is the lag order of ; is a constant vector representing the stable value of the state without any external influence; is the number of samples (total observation time); is a dimension coefficient matrix describing the influence of the past vector on the current vector ; is a
[0128] dimension disturbance vector representing the difference between the actual value and the predicted value caused by random disturbance or model error.
[0129] The BVAR model requires time series to be stationary, so data stationarity test is needed. ADF method is used to test whether the time series data has a unit root, to avoid "pseudo-regression" problem and eliminate trends and seasonality in the data. First-order difference formula is used to convert non-stationary time series into stationary time series. The core formula of ADF test is as follows:
[0130]
[0131] In the formula, is the monitoring value at ; is the first-order difference, representing the daily change of the monitoring value, used to eliminate short-term fluctuations; is the constant term, reflecting the deviation of the sequence mean; is the time trend term, which is included when the sequence has a clear trend; is the coefficient of the lagged difference term, used to eliminate high-order autocorrelation; is the lag order, determined by information criterion; These are core parameters used to determine whether a sequence is stationary; This is the error term.
[0132] When the ADF statistic (core parameter) of When the absolute value of the statistical statistic is less than the critical significance level, it needs to be subjected to first differencing to eliminate the trend and make the series stationary; otherwise, it can be directly used for model calculation. The first differencing formula is as follows:
[0133]
[0134] Perform the ADF test again on the differencing sequence until the sequence becomes stationary. Generally, a single first-order differencing is sufficient to meet the stationarity requirement.
[0135] Step S23: Determine the optimal lag order using the BIC information criterion.
[0136] The information criterion seeks a balance between model goodness of fit and model simplicity. Its goal is to select the best model from all candidate models. The model with the smallest value, that is, the model with the largest posterior probability, is the optimal model.
[0137]
[0138] In the formula, The maximum likelihood estimate of the model represents the goodness of fit of the model to the data. The number of parameters (degrees of freedom) in a model reflects its complexity. Given the sample size, the number of parameters increases accordingly. and sample size The increase, The penalties are also increasing.
[0139] Step S24: Calculate the hysteresis parameters using the impulse response function (IRF):
[0140] Time-domain impulse response analysis This model quantifies the intensity and direction of the impact of a unit shock on other variables at different lag times in a BVAR model, revealing the dynamic interactions between variables. It describes the series of effects on other variables (such as seepage) over a current and future period when a variable (e.g., reservoir water level) is subjected to a unit-sized "shock" (change). The impulse response function can be expressed as:
[0141]
[0142] In the formula, The response value of the affected variable; For the affected variables; For the source of the impact; for exist The unit impact applied at any given time; For response time.
[0143] The optimal lag order of the model was determined to be 6 using the BIC information criterion, and the lag days of reservoir water level were obtained. Days, affecting the number of days Day; Temperature hysteresis parameter sky, sky.
[0144] like Figure 2 As shown, the reservoir water level consistently exhibits a positive impulse response intensity for seepage, with a relatively long response duration. This indicates that changes in reservoir water level have a continuous and stable positive driving effect on seepage. Specifically, when the reservoir water level rises, it significantly promotes an increase in seepage flow through physical mechanisms such as increased pore water pressure and enhanced permeability gradient; while when the reservoir water level falls, this positive effect weakens somewhat, but still maintains a relatively long-lasting effect. The influence of temperature on seepage exhibits more complex characteristics: the impulse response intensity fluctuates within the positive and negative range, and the overall response intensity is relatively weak, with a shorter duration. On the one hand, increased temperature may promote seepage by reducing water viscosity and enhancing the mobility of water molecules; on the other hand, the thermal expansion and contraction effect of rock mass caused by temperature changes may alter fracture aperture, thereby inhibiting seepage.
[0145] Step S3: Feature factor importance screening:
[0146] Step S31, Factor Contribution Analysis:
[0147] Too many influencing factors can lead to feature redundancy and cause overfitting in prediction accuracy. By selecting factors, the number of factors can be reduced, thereby improving the interpretability and generalization ability of the model. The algorithm introduces the concept of weight calculation, and achieves feature selection by considering the weight differences between neighboring samples of different categories. The specific steps of the algorithm are as follows:
[0148] 1) Initialize the number of iterations and the number of nearest neighbors selected .
[0149] 2) Randomly select objects from the dataset From and Select from similar objects The nearest neighbor is denoted as ,and Selecting from heterogeneous objects The nearest neighbor is denoted as , repeat until the iteration sampling number is .
[0150] 3) Calculate the distance between 2 objects on the feature :
[0151]
[0152] where, and are the values of the objects and on the feature ; and are the maximum and minimum values of the feature .
[0153] The importance value :
[0154]
[0155] where, is the iteration number, ; is the probability of the object belonging to the same category ; is the probability of the object belonging to the different category .
[0156] 4) Repeat the above steps to output the importance vector , then the feature weight corresponding to the feature is:
[0157]
[0158] Step S32, VIF collinearity test:
[0159] When there is serious multicollinearity between factors, the independent variables between each other affect and change, so that the contribution of the independent variable to explain the ability is weakened, thereby affecting the accuracy of the prediction result. The variance inflation coefficient (VIF) is used to test the multicollinearity of the factors, and the calculation formula is as follows:
[0160]
[0161] In the formula: represents the fitting coefficient of the independent variable.
[0162] The multiple collinearity of the variable is judged by the size of the VIF value, and the greater the VIF value, the stronger the collinearity between factors.
[0163] Step S33, characteristic factor screening:
[0164] According to the results of steps S31 and S32, in combination with the importance ranking and comprehensive screening, the factors with low importance in the high VIF factors are removed, and finally the optimized characteristic factor set with high prediction performance, strong interpretability and good generalization ability is determined, thereby providing reliable characteristic input for constructing a high-precision underground powerhouse seepage prediction model.
[0165] The application further provides an underground powerhouse seepage characteristic factor set construction system considering hysteresis effect, which comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor, and when the processor executes the computer program instructions, the method described above can be realized.
[0166] The application further provides a computer readable storage medium, which stores computer program instructions capable of being executed by a processor, and when the processor executes the computer program instructions, the method described above can be realized.
[0167] The above is the preferred embodiment of the application, and any change made according to the technical solution of the application, as long as the function does not exceed the scope of the technical solution of the application, belongs to the protection scope of the application.
Claims
1. A method for constructing a set of seepage characteristic factors of an underground powerhouse considering hysteresis effects, characterized in that, include: Step S1: Construction of the feature factor set for the seepage prediction model of the underground powerhouse of the pumped storage power station: By analyzing the seepage mechanism and influencing factors of the underground powerhouse of the pumped storage power station, an initial set of characteristic factors covering reservoir pressure, temperature, rainfall and time factors was obtained; Step S2: Solving for the lag characteristic factor parameters: By establishing a Bayesian vector autoregression model and applying impulse response function analysis, the lag days and influence days of factors including outflow water level and temperature on seepage are solved, and the lag characteristic factors are obtained. Step S3: Feature factor importance screening: Factor contribution analysis was performed using the ReliefF algorithm, collinearity test was conducted using the variance inflation coefficient, and factor importance ranking and comprehensive screening were combined to finally determine an optimized set of feature factors that have high predictive performance, strong interpretability and good generalization ability, providing reliable feature input for building a high-precision underground powerhouse seepage prediction model.
2. The method of claim 1, wherein the method further comprises: In step S1, the initial feature factor set is represented as: wherein, , , , respectively represent the water pressure component, the temperature component, the rainfall component and the time component.
3. The method of claim 1, wherein the method further comprises: Step S1 is as follows: Step S11, Monitoring data preprocessing: The raw seepage data of the underground powerhouse of the pumped storage power station was preprocessed to repair missing and outlier values. At the same time, the data from different sources were standardized and multi-source heterogeneous data were combined into a unified format. Step S12, Construction of the seepage characteristic factor set: The influencing factors of seepage effect in the underground powerhouse of pumped storage power station were analyzed, and an initial feature factor set was constructed for the seepage prediction model of the underground powerhouse of pumped storage power station.
4. The method of claim 3, wherein the method further comprises: In step S11, the missing value is handled using cubic spline interpolation.
5. The method of claim 3, wherein the method further comprises: In step S11, the selection of the monitoring data sampling area should be based on the distribution of seepage monitoring points in the underground powerhouse of the pumped storage power station, engineering geological conditions and potential seepage paths, and should be determined comprehensively by combining the existing monitoring points and actual needs.
6. The method of claim 3, wherein the method further comprises: In step S11, the sampling time should cover at least one complete hydrological year and include the characteristic operating conditions of the reservoir during its operation: flood season, discharge season, and dry season.
7. The method of claim 1, wherein the method further comprises: Step S2 is as follows: Step S21: Establish a BVAR Bayesian vector autoregressive model: The BVAR Bayesian vector autoregressive model is represented as follows: In the formula: for A 3D state vector represents the state of different variables at different times. The value at time, The number of variables; for The lag order; is a constant vector representing the stable values of the state when there are no external influences; The number of samples; for A coefficient matrix describing the past vector at time points For the current vector The impact; for A perturbation vector, representing the difference between the actual and predicted values caused by random perturbations or model errors; Step S22, ADF data stationarity test: The ADF test is used to examine whether a unit root exists in time series data, avoiding the "spurious regression" problem and eliminating trends and seasonality in the data. The first-order difference formula is used to transform non-stationary time series into stationary ones. The core formula of the ADF test is as follows: In the formula, is the monitoring value at the moment; is the first-order difference, representing the daily variation of the monitoring value, and used to eliminate short-term fluctuations; is the constant term, reflecting the deviation of the sequence mean; is the time trend term; is the coefficient of the lagged difference term, used to eliminate high-order autocorrelation; is the lag order, determined by the information criterion; is the core parameter, used to determine whether the sequence is stationary; is the error term; When the ADF statistic is of When the absolute value of the statistic is less than the critical significance level, it indicates that the series is in a non-stationary state and needs to be first-differenced to eliminate the trend and make the series stationary. The ADF test is then performed on the differencing series again until the series becomes stationary. Step S23: Determine the optimal lag order using the BIC information criterion. The information criterion aims to select the model with the smallest value, i.e. the model with the largest posterior probability, among all candidate models: The information criterion aims to select the model with the smallest value, i.e. the model with the largest posterior probability, among all candidate models: wherein is the maximum likelihood estimate of the model, indicating the goodness of fit of the model to the data; is the number of parameters of the model, reflecting the complexity of the model; is the sample size, as the number of parameters and the sample size increase, the penalty term also increases; Step S24: Calculate the hysteresis parameters using the impulse response function (IRF): The impulse response function (IRF) quantifies the strength and direction of the impact of a unit shock on other variables in a Bayesian vector autoregressive (BVAR) model at different lag times. It describes the series of effects on other variables over a period of time, both now and in the future, when a variable is subjected to a unit-sized "shock" (change). The IRF is expressed as: wherein, is a response value of the affected variable; is an affected variable; is an impact source variable; is at a unit impact applied at a time instant; is a response time; The lag days and impact days obtained from the impulse response function (IRF) analysis are used as key parameters and substituted into the lag effect function to obtain the lag characteristic factor.
8. The method of claim 1, wherein the method further comprises: Step S3 is as follows: Step S31, Factor Contribution Analysis: The algorithm introduces the concept of weight calculation, and realizes the feature selection function by considering the weight difference between different categories of neighbor samples, The specific steps of the algorithm are as follows: (1) initialize the number of iterations with the selected number of neighbors ; (2) Randomly select objects from the dataset From and Select from similar objects The nearest neighbor is denoted as ,and Selecting from heterogeneous objects The nearest neighbor is denoted as Repeat this process until the number of iterations is reached. ; (3) Calculate the distance between 2 objects In the feature on the distance is: wherein and are the values of the object and on the feature ; and are the maximum and minimum values of the feature , respectively. importance value : wherein is the number of iterations, ; is the probability that the object belongs to the same class ; is the probability that the object belongs to a different class ; (4) Repeat steps (1)-(3) to output the importance vector Then the feature The corresponding feature weight Is: Step S32, VIF collinearity test: The multiple collinearity of factors is tested by using the variance inflation factor VIF, and the calculation formula is as follows: wherein: represents the fitting coefficient of the independent variable; The multiple collinearity of variables is judged by the size of VIF value, and the larger the VIF value, the stronger the collinearity between factors; Step S33, characteristic factor screening: According to the results of steps S31 and S32, the factors with low importance in the high VIF factors are removed by combining the factor importance ranking and comprehensive screening, and finally the optimized characteristic factor set with high prediction performance, strong interpretability and good generalization ability is determined, which provides reliable feature input for constructing high-precision underground powerhouse seepage prediction model.
9. A system for constructing a set of seepage characteristic factors of an underground powerhouse considering hysteresis effects, characterized in that, The computer program instructions stored on the memory and capable of being run by the processor can realize the underground powerhouse seepage characteristic factor set construction method considering the hysteresis effect according to any one of claims 1-8 when the processor runs the computer program instructions.
10. A computer-readable storage medium, characterized in that, The computer program instructions stored on the memory and capable of being run by the processor can realize the underground powerhouse seepage characteristic factor set construction method considering the hysteresis effect according to any one of claims 1-8 when the processor runs the computer program instructions.