A method for hierarchical screening and recursive optimization of key driving factors for medium and long term runoff prediction

CN122840318APending Publication Date: 2026-09-29STATE GRID CORP NORTHEAST DIVISION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610886961.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明要解决的技术问题是针对现有中长期径流预测中因子数量多、冗余性强、时滞关系复杂以及关键驱动信息难以稳定识别的问题,提出一种面向中长期径流预测的关键驱动因子分层筛选与递归优化方法,是一种能够实现多源驱动因子冗余消减、独立信息增益识别与关键驱动因子动态筛选的方法,以提高中长期径流预测模型的泛化能力、稳定性与可解释性

Benefits of technology

[0022]本发明突破了传统相关系数法、逐步回归法等单因子筛选方法仅关注因子与径流之间线性相关关系的局限,通过构建基于RV相关系数的相似因子划分机制,能够有效压缩同质化信息;构建的因子族内部竞争筛选机制,通过统一预测模型和交叉验证评价体系对同一信息因子族内的候选因子进行性能比较,能够从多个具有相似信息特征的因子中选取代表性最强的驱动因子,提高筛选效率并降低特征维度;进一步引入的增量贡献分析方法,通过定量评估新增因子对模型预测性能的边际贡献,保证保留因子具有独立有效的信息增益;基于沙普利加性解释方法构建的递归淘汰机制,能够从全局解释角度量化各预测变量对径流预测结果的贡献程度,并依据贡献度逐步优化预测变量组合,提高最终关键驱动因子集合的可解释性和预测有效性。本发明形成的一套兼顾信息冗余控制、增量信息识别和模型可解释性的关键驱动因子筛选方法,可有效解决中长期径流预测中候选因子数量庞大、信息重叠严重、筛选结果不稳定以及物理解释不足等问题,为构建高精度、高可靠性的中长期径流预测模型提供技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840318A_ABST
    Figure CN122840318A_ABST
Patent Text Reader

Abstract

The application belongs to the field of hydrological prediction, and particularly relates to a key driving factor hierarchical screening and recursive optimization method for medium and long term runoff prediction. The application is based on multi-source driving data such as teleconnection factors, astronomical factors and hydrological and meteorological factors of a research basin, constructs a factor family containing multi-time lag characteristics as a prediction variable; selects RV correlation quantity to quantify the information similarity between different factor families, and adopts a hierarchical clustering method to realize similar factor merging; performs prediction based on a LightGBM regression model, constructs an internal competition mechanism of a clustering cluster, identifies factors with independent information gain through single factor prediction evaluation and incremental contribution analysis; finally, a Shapley additive explanation method is introduced to recursively screen the prediction variables, gradually eliminate prediction variables with low contribution degree, and obtain input variables corresponding to an optimal prediction performance model as a key driving factor set for medium and long term runoff prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrological forecasting, specifically involving a hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting. Background Technology

[0002] Medium- and long-term runoff forecasting is of great guiding significance for regional water resource allocation, flood control and disaster reduction, and reservoir operation (Xu Zongxue, Zhou Zuhao, Jiang Yao, et al. Variation law of runoff in the source areas of rivers in Southwest China and its future evolution trend [J]. Advances in Water Science, 2022, 33(3):360-374.). However, due to the multi-dimensional nonlinear characteristics of hydrological systems, the accuracy of medium- and long-term runoff forecasting is a major challenge for the industry, especially as the forecast period becomes longer, the forecast accuracy is usually further weakened. Improving the accuracy of medium- and long-term runoff forecasting can be started from the selection of forecasting factors. In recent years, the introduction of teleconnected climate and astronomical factors into traditional predictive factors has been proven to effectively improve the accuracy of medium- and long-term runoff prediction (Li Jiqing, Xie Yutao, Xu Xuejun, et al. Teleconnected prediction of monthly runoff based on RF-Informer model [J]. Water Resources Protection, 2025, 41(3):39-45.). However, due to the large number of categories of teleconnected and astronomical factors, the complex relationships between factors, and the varying degrees of influence on runoff, extracting key driving factors can not only improve the training speed of the prediction model, but also eliminate weakly correlated factors, reduce noise, and further improve prediction accuracy.

[0003] Currently, medium- and long-term runoff prediction models face challenges such as high factor dimensionality, complex time-lag relationships, and significant multi-factor coupling. Many factors often exhibit strong correlations and information overlap. Traditional factor selection methods typically rely solely on correlation coefficients, univariate statistical indicators, or single-model importance, making it difficult to distinguish the true independent contributions of different factors. This can easily lead to a large influx of redundant factors into the model, resulting in overfitting, decreased prediction stability, and insufficient generalization ability (Chen Juan, Xu Qi, Cao Duanxiang, et al. Medium- and long-term runoff prediction model based on multi-factor and multi-model integration [J]. Advances in Water Science, 2024, 35(3):408-419.). Furthermore, medium- and long-term runoff processes exhibit significant non-stationarity and multi-timescale response characteristics, with the effects of different driving factors dynamically changing at different time stages. Traditional methods lack the comprehensive analytical capabilities for factor increment information, dynamic contribution capabilities, and multi-timescale lag effects, making it difficult to reliably identify key driving factors that truly have sustained predictive value for future runoff changes. Summary of the Invention

[0004] The technical problem this invention aims to solve is the issue of numerous factors, high redundancy, complex time-delay relationships, and difficulty in stably identifying key driving information in existing medium- and long-term runoff forecasting. It proposes a hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting. This method can achieve redundancy reduction of multi-source driving factors, identification of independent information gain, and dynamic screening of key driving factors, thereby improving the generalization ability, stability, and interpretability of medium- and long-term runoff forecasting models.

[0005] This invention constructs a similar factor family division rule based on RV correlation coefficient, and combines internal competition within factor families, incremental contribution analysis, and a recursive elimination mechanism based on Shapley additive interpretation method. By jointly analyzing the redundancy relationship, independent information gain, and dynamic contribution capability of multi-source driving factors, it achieves stable identification and optimized screening of key driving factors for medium- and long-term runoff prediction, thereby improving the generalization ability, stability, and interpretability of the prediction model.

[0006] The technical solution of this invention is as follows:

[0007] A hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting includes the following steps:

[0008] Step (1) Data collection and preprocessing

[0009] A candidate factor library was formed by collecting historical monthly runoff data of the watershed and a dataset of candidate driving factors affecting runoff changes. These candidate driving factors include teleconnection factors, astronomical factors, and watershed-scale hydro-meteorological factors. After unifying the time scale and handling missing values ​​for the original sequences of each factor, the starting month of the prediction model was set as t, the prediction period as Y months, and the target month as t+Y. For any factor, based on the starting time t, a time-lag feature sequence was constructed consisting of the corresponding months t, t-1, t-2, ..., t-11, serving as a candidate predictor variable for the target month t. The time-lag feature sequences formed by the same factor constitute the factor family for that factor, as illustrated in Table 1.

[0010] Table 1. Schematic diagram of the correspondence between prediction targets and prediction variables.

[0011]

[0012] Step (2) Construct the factor similarity matrix

[0013] Calculate the RV correlation coefficient between any two candidate factors, and construct a candidate factor similarity matrix using the RV correlation coefficient to quantify the degree of information consistency between different candidate factors and characterize the similarity relationship between candidate factors.

[0014] Step (3) Information similarity factor family division

[0015] Based on the factor similarity matrix, a hierarchical clustering algorithm is used to cluster candidate factors; factors carrying similar information features are grouped into the same cluster, thereby merging redundant information of candidate factors.

[0016] Step (4) Intra-cluster competitive screening

[0017] For each candidate factor in each cluster, a single-factor prediction model based on the LightGBM regression model is constructed. In one implementation, the LightGBM regression model is implemented using the LGBMRegressor regressor provided by the lightgbm library in the Python environment. Specifically, a training sample set is constructed using the time lag feature sequences of each candidate factor as the model input variable and the measured runoff of the corresponding target month as the target value for model training. A time-sequence-based cross-validation method is used to train and validate the training sample set in multiple rounds. In each round of validation, the training subset is input into the LightGBM regression model for fitting training, and the validation subset is input into the trained model, outputting the runoff prediction value for the corresponding target month. Based on the validation set determination coefficients obtained in each round of validation, the predictive ability and generalization performance of the corresponding candidate factors are evaluated, and the candidate factors within the same cluster are ranked accordingly. The candidate factor with the best predictive performance is selected as the benchmark factor for that cluster. Other candidate factors within the cluster are then combined with the benchmark factor to form a set of predictive variables, thereby constructing a joint prediction model. The performance increment of the prediction model brought by the newly added factors is calculated. New and redundant information is identified based on the magnitude of the increment contribution, and factors that cannot provide effective increment information are eliminated. Finally, factors with non-negative increment contributions from all clusters are summarized to form a set of candidate key driving factors.

[0018] Step (5) Recursive screening based on Shapley's additive interpretation method

[0019] All time-lag feature sequences composed of candidate key driving factors are input as predictor variables into the LightGBM regression model described in step (4) for training. After training, the contribution of each predictor variable to the model prediction result is calculated using the Shapley additive interpretation method; wherein, the contribution is the average absolute Shapley value of each predictor variable in the validation sample, used to characterize the influence of the predictor variable on the model prediction result. The predictor variables are sorted from smallest to largest according to their contribution, and the predictor variable with the lowest contribution is removed in each iteration; after each removal, the runoff prediction model is retrained using the remaining predictor variables, and the prediction performance of the retrained model is evaluated according to the preset model performance evaluation index. The process of contribution calculation, predictor variable removal, model retraining, and performance evaluation is repeated until only one predictor variable remains. The set of predictor variables corresponding to the runoff prediction model with the best prediction performance in the iteration process is determined as the set of key driving factors for long-term runoff prediction in the study watershed.

[0020] The present invention also provides a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the above-described method for screening key driving factors of medium- and long-term runoff.

[0021] The beneficial effects of this invention are:

[0022] This invention overcomes the limitations of traditional single-factor screening methods such as correlation coefficient method and stepwise regression method, which only focus on the linear correlation between factors and runoff. By constructing a similar factor partitioning mechanism based on RV correlation coefficient, it can effectively compress homogeneous information. The constructed factor family internal competition screening mechanism compares the performance of candidate factors within the same information factor family through a unified prediction model and cross-validation evaluation system, which can select the most representative driving factor from multiple factors with similar information characteristics, improving screening efficiency and reducing feature dimensionality. The further introduced incremental contribution analysis method quantitatively evaluates the marginal contribution of newly added factors to the model's predictive performance, ensuring that the retained factors have independent and effective information gain. The recursive elimination mechanism constructed based on Shapley additive interpretation method can quantify the contribution of each predictor variable to the runoff prediction results from a global interpretation perspective, and gradually optimize the combination of predictor variables according to the contribution, improving the interpretability and predictive effectiveness of the final key driving factor set. The present invention provides a key driver factor screening method that takes into account information redundancy control, incremental information identification and model interpretability. It can effectively solve the problems of large number of candidate factors, serious information overlap, unstable screening results and insufficient physical interpretation in medium and long-term runoff prediction, and provide technical support for building high-precision and high-reliability medium and long-term runoff prediction models. Attached Figure Description

[0023] Figure 1 A flowchart illustrating the method of the present invention.

[0024] Figure 2 It is a hierarchical clustering tree of factor family structure similarity constructed based on RV correlation coefficient. Detailed Implementation

[0025] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0026] A hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting, the process is as follows: Figure 1 As shown, the specific steps are as follows:

[0027] Step (1) Data collection and preprocessing;

[0028] 1.1) Obtain historical data of teleconnection factors, historical data of astronomical factors, and watershed hydro-meteorological factors for the target watershed.

[0029] The target watershed in this study is the Fengman Reservoir within a specific watershed. Teleconnection factors include 130 large-scale meteorological-climate indices, comprising 88 monthly atmospheric circulation indices, 26 monthly sea surface temperature indices, and 16 other indices. Data are from the National Climate Center of the China Meteorological Administration (https: / / cmdp.ncc-cma.net / Monitoring / cn_index_130.php), with a time series from 1951 to 2021. Astronomical factors include lunar declination angle and sunspot number, with data from NASA's Jet Propulsion Laboratory (ssd.jpl.nasa.gov / horizons) and the Royal Observatory of Belgium's WDC-SIL, respectively. The data files (sidc.be / SILSO / datafiles) contain 40 watershed-scale hydrometeorological factors, including historical runoff data, local temperature index, snowfall index, soil water index, and surface radiation index. These factors were obtained from the Northeast China Power Grid and the ERA5 dataset (https: / / cds.climate.copernicus.eu / datasets / reanalysis-era5-land-monthly-means?tab=overview), with the time series spanning from 1951 to 2021.

[0030] 1.2) Data Preprocessing

[0031] Factors that have been missing for more than 12 consecutive months in the above data sequence were removed, and the missing values ​​were then processed by linear interpolation.

[0032] Based on the above data, let the monthly inflow sequence of the target reservoir be... The unit is , Indicates the sequence number of the time series. The total number of months in the time series; candidate factor sequences are ,in This represents the candidate factor index. Considering the time-lag response effect of driving factors to runoff changes, in this example, with a forecast period of Y months, the starting month index of the prediction model is t, and the prediction target is expressed as... The time-lag feature sequence formed by each candidate factor constitutes The set of predictor variables at time t is represented as:

[0033] (1)

[0034] Therefore, the multi-time-delay mapping relationship between the driving factors and the target runoff can be expressed as:

[0035] (2)

[0036] in, This represents the functional mapping relationship between predictor variables and target runoff.

[0037] Step (2) Construct the factor similarity matrix

[0038] Original sequences of each candidate factor Perform time delay processing to construct a multi-delay matrix:

[0039] (3)

[0040] in, Indicates the first A multi-delay matrix of candidate factors.

[0041] Let any number of... , The multi-delay matrix of the two factors is , The mean of each column in each matrix is ​​removed:

[0042] (4)

[0043] in:

[0044] (5)

[0045] in, , They represent , The mean vector of each column; , These represent the multi-delay matrices after mean removal processing; This represents the sequence number in a time series, with a step size of months. This indicates transpose.

[0046] Calculate the between-sample inner product matrix:

[0047] (6)

[0048] in, , They represent the factor families respectively , The structural relationships between the samples described.

[0049] Calculate the RV correlation coefficient:

[0050] (7)

[0051] in, This represents the trace of a matrix, which is the sum of the elements on the main diagonal. Indicates the first , The RV correlation coefficient values ​​among the factors The larger the RV correlation coefficient, the closer the information structures carried by the two factor families; the smaller the RV correlation coefficient, the greater the information differences between the two factor families. Similarly, calculate the RV correlation coefficients between other factors pairwise. Therefore, the factor similarity matrix... It can be represented as:

[0052] (8)

[0053] Step (3) Information similarity factor family division

[0054] 3.1) Construct the factor family distance matrix

[0055] Since hierarchical clustering algorithms typically use a distance matrix as input, and the RV correlation coefficient is a similarity metric, it is necessary to convert the similarity matrix into a distance matrix. This invention employs a linear transformation method, defining the distance between factor families as:

[0056] (9)

[0057] in, Representation by family With factor family The distance between them.

[0058] Therefore, the distance matrix of all factor families can be expressed as:

[0059] (10)

[0060] 3.2) Information similarity factor family division based on hierarchical clustering

[0061] Hierarchical clustering is used to perform cluster analysis on the factor family distance matrix. To meet the input requirements of the hierarchical clustering algorithm, the factor family distance matrix is ​​converted into a compressed distance vector form. The data is used as input to the hierarchical clustering algorithm. Based on compressed distance vectors, the average connectivity method is used to perform hierarchical clustering analysis. In the initial stage of clustering, each factor family is treated as an independent cluster. During the clustering process, the two clusters with the smallest current distance are continuously searched and merged. After each merge, the distance relationship between the new cluster and the remaining clusters is recalculated until all factor families finally form a complete hierarchical structure. The result is plotted as a clustering tree and the distance threshold is used to define the clustering. (In this embodiment, the value is 0.3) The clustering tree is cut, and the clustering tree is divided into segments based on distance. The merged factor families are grouped into the same cluster, completing the information similarity factor family partitioning; see the clustering tree for this example. Figure 2 .

[0062] During the clustering process, information about each cluster merge is recorded, including the two clusters being merged, the distance value at the time of the merge, and the number of factor families contained in the merged cluster. If two factor families merge at a low distance, it indicates that their RV structure is highly similar and the information features they carry are relatively close; if two factor families merge at a high distance, it indicates that the information they reflect is significantly different.

[0063] Step (4) Intra-cluster competitive screening

[0064] 4.1) Selection of optimal factors for clustering

[0065] Let the first Each cluster is represented as:

[0066] (11)

[0067] in, Indicates the first The first cluster in the cluster One candidate factor, , This represents the number of factors in the cluster.

[0068] For any factor Extract all the corresponding lagged features, that is, use the time lagged feature sequence of the same factor as the predictor variable to input into the prediction model.

[0069] With target runoff sequence As the prediction target, a LightGBM regression model was used to establish a nonlinear mapping relationship between the predictor variables and runoff. The coefficients of determination for the training set and the validation set were obtained through 5-fold time series cross-validation. Subsequently, the LightGBM regression model was trained using all samples, and the root mean square error (RMSE) was calculated after obtaining the corresponding prediction results. RMSE measures the deviation between the predicted results and the actual observed values; the smaller the RMSE value, the smaller the model fitting error and the higher the prediction accuracy. The calculation formula is as follows:

[0070] (12)

[0071] in, Indicates the sequence number of the time series; This represents the actual runoff value. To predict runoff values.

[0072] For all candidate factors in the same cluster, they are sorted in descending order according to the validation set determination coefficient, and the factor with the highest score is selected as the benchmark factor for that cluster.

[0073] 4.2) Incremental Contribution Analysis

[0074] After the intra-cluster competitive screening in step 4.1), each cluster obtains a baseline factor. Since other factors within the same cluster may still carry some supplementary information, it is necessary to further evaluate the contribution of each candidate factor to the incremental information relative to the baseline factor, thereby identifying redundant and incremental information. To this end, this invention constructs an incremental contribution analysis mechanism based on prediction performance gain, which quantitatively evaluates the independent information value of candidate factors by comparing the changes in model prediction performance before and after the addition of new factors.

[0075] First, the competition results within each cluster are pre-screened, retaining only factors with a validation set determination coefficient of not less than zero. The LightGBM regression model described in step 4.1) is used as the prediction model, with the baseline factor and the prediction variables corresponding to other factors in the same cluster serving as inputs. The model output is the predicted runoff value for the target time period. A 5-fold time series cross-validation method is used to train and validate the model. The determination coefficient is calculated based on the predicted and measured runoff values ​​in the validation set. The incremental contribution of a new factor is defined as the difference between the validation set determination coefficient after introducing the factor and the validation set determination coefficient when only the baseline factor is used. A larger incremental contribution value indicates higher independent information value for the factor. Finally, the baseline factor and factors with incremental contribution values ​​greater than 0 are retained, while other candidate factors are eliminated, forming a set of candidate key driving factors.

[0076] Step (5) Recursive screening based on Shapley's additive interpretation method

[0077] This invention further introduces a recursive factor selection mechanism based on the Shapley additive interpretation method. By iteratively evaluating the contribution of factors to the model prediction results, the factors with the weakest contribution are gradually eliminated, thereby obtaining the optimal factor combination that balances prediction performance and model simplicity.

[0078] First, based on the incremental contribution analysis obtained in step (4), the set of candidate key driving factors is obtained, and all their corresponding time lag features are extracted as the initial set of predictive variables for recursive screening. Then, the LightGBM regression model obtained in the above steps is used for training and validation using the 5-fold time series cross-validation method. The coefficient of determination of the validation set is calculated and used as one of the performance evaluation indicators of the model. Subsequently, the LightGBM regression prediction model is retrained using all samples to obtain the runoff prediction value. The root mean square error (RMSE) and mean absolute error (MAE) are calculated. RMSE and MAE are used to evaluate the overall fitting error level of the model. The formula for calculating MAE is as follows:

[0079] (13)

[0080] After model training is complete, the contribution of each input feature to the model prediction process is quantitatively analyzed based on the Shapley additive interpretation method. The absolute value of the Shapley value represents the absolute magnitude of the feature's contribution to the model output. The Shapley value of each feature is calculated using the following formula:

[0081] (14)

[0082] In the formula: For the first The Shapley value of each feature; M is the number of features; ! represents factorial; For the input feature set, ; For features not included Input feature set; For input The model used for forecasting is the LightGBM regression prediction model mentioned above; For features not included A subset of features; For set The number of features.

[0083] Based on formula (14), the Shapley value of each predictor variable in each sample is calculated, and the mean absolute Shapley value of each predictor variable is calculated. The larger the mean absolute Shapley value, the more significant the contribution of the predictor variable to the runoff prediction result. All predictor variables are sorted in ascending order according to their contribution. The predictor variable with the smallest contribution is eliminated from the current predictor variable set and then re-input into the LightGBM regression model. Time series cross-validation, full sample training, and Shapley interpretation are performed. Only one predictor variable is deleted in each iteration. During the recursive elimination process, the model performance evaluation results corresponding to each iteration are recorded. The largest coefficient of determination in the validation set is used as the criterion for the optimal model. The root mean square error (RMSE) and mean absolute error (MAE) are used as auxiliary evaluation indicators to determine the predictor variable set corresponding to the optimal model. This set is used as the key driving factor set for medium- and long-term runoff prediction.

[0084] Taking Fengman Reservoir as the research object, 172 candidate influencing factors were collected. After data preprocessing, 155 effective factors were retained, and 1860 predictive variables were constructed through lag expansion. Following the method described in this invention, 53 key predictive variables were finally obtained, constituting the optimal input feature set. Furthermore, the LightGBM prediction models involved in this invention all adopt a unified model structure and hyperparameter configuration, the specific parameter settings of which are shown in Table 2.

[0085] Table 2 Hyperparameter settings for the LightGBM model

[0086]

[0087] To verify the effectiveness of the multi-stage feature screening method proposed in this invention, three sets of comparative experiments were constructed: (1) Scheme 1: all 1860 original predictor variables were used as model inputs; (2) Scheme 2: predictor variables selected based on Kendall correlation coefficients were used as model inputs; (3) Scheme 3: 53 key predictor variables selected in this invention were used as model inputs. Under the condition that the LightGBM regression model structure, hyperparameter configuration and training strategy were kept completely consistent, the monthly runoff of Fengman Reservoir was predicted, and the prediction performance of each model was compared and analyzed. The results are shown in Table 3.

[0088] Table 3 Comparison of model prediction performance under different strategies

[0089]

[0090] As shown in Table 3, compared with Scheme 1 and Scheme 2, the method proposed in this invention improves the validation set determination coefficient by 48% and 132% respectively, reduces RMSE by 85% and 86% respectively, and reduces MAE by 82% for both. This indicates that the present invention can effectively eliminate redundant information and noise interference in high-dimensional features, extract the key variables that contribute most to runoff prediction, and improve the prediction accuracy and stability of the model while significantly reducing model complexity, thus having good engineering application value.

[0091] This application also provides a computer device, specifically a computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0092] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0093] In one embodiment, a computer-readable storage medium is also provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0094] In one embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0096] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0097] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0098] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0099] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.

Claims

1. A hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting, characterized in that, Includes the following steps: Step (1) Data collection and preprocessing A candidate factor library was formed by collecting and studying historical monthly runoff data of the watershed and a dataset of candidate driving factors affecting runoff changes. The candidate driving factors include teleconnection factors, astronomical factors, and watershed-scale hydrometeorological factors. After unifying the time scale and handling missing values ​​for the original sequence of each factor, let the starting month number of the prediction model be t, the forecast period be Y months, and the target month be t+Y. For any factor, based on the reporting start time t month, a time lag feature sequence is constructed consisting of the corresponding months t, t-1, t-2, ..., t-11 of the factor, which serves as a candidate predictor variable for the target month t. The time lag feature sequences formed by the same factor constitute the factor family of the factor. Step (2) Construct the factor similarity matrix Calculate the RV correlation coefficient between any two candidate factors, and construct a candidate factor similarity matrix using the RV correlation coefficient to quantify the degree of information consistency between different candidate factors and characterize the similarity relationship between candidate factors; Step (3) Information similarity factor family division Based on the factor similarity matrix, a hierarchical clustering algorithm is used to cluster candidate factors; factors carrying similar information features are grouped into the same cluster, thereby merging redundant information of candidate factors; Step (4) Intra-cluster competitive screening For each candidate factor in each cluster, a single-factor prediction model based on the LightGBM regression model is constructed. Specifically, the time lag feature sequence of each candidate factor is used as the model input variable, and the measured runoff of the corresponding target month is used as the target value for model training to construct a training sample set. The training sample set is trained and validated in multiple rounds using a cross-validation method based on time order partitioning. In each round of validation, the training subset is input into the LightGBM regression model for fitting training, and the validation subset is input into the trained model to output the runoff prediction value for the corresponding target month. Based on the validation set determination coefficients obtained in each round of validation, the predictive power and generalization performance of the corresponding candidate factors are evaluated, and the candidate factors within the same cluster are ranked accordingly. The candidate factor with the best predictive performance is selected as the benchmark factor for that cluster. The other candidate factors within the cluster are combined with the baseline factor to form a set of predictive variables, thereby constructing a joint prediction model. The performance increment of the prediction model brought by the new factor is calculated. New and redundant information are identified according to the magnitude of the increment contribution. Factors that cannot provide effective increment information are eliminated. Finally, all factors with non-negative increment contributions in all clusters are summarized to form a set of candidate key driving factors. Step (5) Recursive screening based on Shapley's additive interpretation method All time-lag feature sequences composed of candidate key driving factors are input as predictor variables into the LightGBM regression model described in step (4) for training; after training, the contribution of each predictor variable to the model prediction result is calculated using the Shapley additive interpretation method; wherein, the contribution is the average absolute Shapley value of each predictor variable in the validation sample, which is used to characterize the influence of the predictor variable on the model prediction result; the predictor variables are sorted from smallest to largest according to their contribution, and the predictor variable with the lowest contribution is removed in each iteration; after each removal, the runoff prediction model is retrained using the remaining predictor variables, and the prediction performance of the retrained model is evaluated according to the preset model performance evaluation index; the process of contribution calculation, predictor variable removal, model retraining and performance evaluation is repeated until only one predictor variable remains; The set of predictive variables corresponding to the runoff prediction model with the best prediction performance during the iteration process is determined as the set of key driving factors for long-term runoff prediction in the study watershed.

2. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 1, characterized in that, Step (1) is as follows: 1.1) Obtain historical data of teleconnection factors, historical data of astronomical factors, and watershed hydro-meteorological factors for the target watershed. 1.2) Data Preprocessing Factors that have been missing for more than 12 consecutive months in the above data sequence were removed, and the missing values ​​were processed by linear interpolation. Based on the above data, let the monthly inflow sequence of the target reservoir be... The unit is , Indicates the sequence number of the time series. This represents the total number of months in the time series. Candidate factor sequence is ,in Let represent the candidate factor index; given a forecast period of Y months, the starting month index of the prediction model is t, and the prediction objective is expressed as... The time-lag feature sequence formed by each candidate factor constitutes The set of predictor variables at time t is represented as: (1) Therefore, the multi-time-delay mapping relationship between the driving factors and the target runoff is established as follows: (2) in, This represents the functional mapping relationship between predictor variables and target runoff.

3. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 1, characterized in that, Step (2) is as follows: Original sequences of each candidate factor Perform time delay processing to construct a multi-delay matrix: (3) in, Indicates the first A multi-delay matrix of candidate factors; Let any number of... , The multi-delay matrix of the two factors is , The mean of each column in each matrix is ​​removed. (4) in: (5) in, , They represent , The mean vector of each column; , These represent the multi-delay matrices after mean removal processing; This represents the sequence number in a time series, with a step size of months. Indicates transpose; Calculate the between-sample inner product matrix: (6) in, , They represent the factor families respectively , The described structural relationships between samples; Calculate the RV correlation coefficient: (7) in, This represents the trace of a matrix, which is the sum of the elements on the main diagonal. Indicates the first , The RV correlation coefficient values ​​among the factors Similarly, calculate the RV correlation coefficients between each pair of other factors; then, the factor similarity matrix... Represented as: (8)。 4. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 1, characterized in that, Step (3) is as follows: 3.1) Construct the factor family distance matrix The similarity matrix is ​​converted into a distance matrix; using a linear transformation method, the distance between factor families is defined as: (9) in, Representation by family With factor family The distance between them; Therefore, the distance matrix for all factor families is expressed as: (10) 3.2) Information similarity factor family division based on hierarchical clustering Hierarchical clustering is used to perform cluster analysis on the factor family distance matrix; to meet the input requirements of the hierarchical clustering algorithm, the factor family distance matrix is ​​converted into a compressed distance vector form. The data is used as input data for the hierarchical clustering algorithm. Based on compressed distance vectors, the average connectivity method is used to perform hierarchical clustering analysis. In the initial stage of clustering, each factor family is regarded as an independent cluster. During the clustering process, the two clusters with the smallest current distance are continuously searched and merged. After each merge, the distance relationship between the new cluster and the remaining clusters is recalculated until all factor families finally form a complete hierarchical structure. The result is plotted as a clustering tree and the distance threshold is used to determine the relationship. The clustering tree is cut, and the clustering tree is divided into segments based on distance. The factor families that have been merged are now grouped into the same cluster, thus completing the information similarity factor family division. During the clustering process, information about each cluster merge is recorded, including the two clusters being merged, the distance value at the time of the merge, and the number of factor families contained in the merged cluster.

5. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 1, characterized in that, Step (4) is as follows: 4.1) Selection of optimal factors for clustering Let the first Each cluster is represented as: (11) in, Indicates the first The first cluster in the cluster One candidate factor, , This represents the number of factors in the cluster. For any factor Extract all the corresponding lagged features, that is, use the time lagged feature sequence of the same factor as the predictor variable and input it into the prediction model. With target runoff sequence As the prediction target, the LightGBM regression model is used to establish a nonlinear mapping relationship between the predictor variables and runoff. The coefficient of determination of the training set and the coefficient of determination of the validation set are obtained by five-fold time series cross-validation. Then, the LightGBM regression model is trained with all samples to obtain the corresponding prediction results. The root mean square error (RMSE) is then calculated. For all candidate factors in the same cluster, they are sorted in descending order according to the coefficient of determination of the validation set, and the factor with the highest score is selected as the benchmark factor of the cluster. 4.2) Incremental Contribution Analysis After the internal competitive screening within the clusters in step 4.1), each cluster obtains a baseline factor; an incremental contribution analysis mechanism based on prediction performance gain is constructed to quantitatively evaluate the independent information value of candidate factors by comparing the changes in model prediction performance before and after the addition of new factors. First, the competition results within the clusters are pre-screened, retaining only factors with a validation set determination coefficient not lower than zero. The LightGBM regression model described in step 4.1) is used as the prediction model, with the baseline factor and the prediction variables corresponding to other factors in the same cluster serving as inputs. The model output is the predicted runoff value for the target time period. A 5-fold time series cross-validation method is used to train and validate the model. The determination coefficient is calculated based on the predicted and measured runoff values ​​in the validation set. The incremental contribution of a new factor is defined as the difference between the model's validation set determination coefficient after introducing the factor and the model's validation set determination coefficient when only the baseline factor is used. Finally, the baseline factor and factors with incremental contribution values ​​greater than 0 are retained, while other candidate factors are eliminated, forming a set of candidate key driving factors.

6. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 5, characterized in that, In step 4.1), the RMSE calculation formula is as follows: (12) in, Indicates the sequence number of the time series; This represents the actual runoff value. To predict runoff values.

7. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 1, characterized in that, Step (5) is as follows: First, based on the incremental contribution analysis obtained in step (4), the candidate key driving factor set is obtained, and all its corresponding time-lag features are extracted as the initial predictive variable set for recursive screening. The LightGBM regression model is used, and the 5-fold time series cross-validation method is adopted for training and validation. The coefficient of determination of the validation set is calculated and used as one of the performance evaluation indicators of the model. Then, the LightGBM regression prediction model is retrained using all samples to obtain the runoff prediction value. The root mean square error (RMSE) and mean absolute error (MAE) are calculated. RMSE and MAE are used to evaluate the overall fitting error level of the model. After the model training is completed, the contribution of each input feature in the model prediction process is quantitatively analyzed based on the Shapley additive interpretation method. The absolute value of the Shapley value represents the absolute magnitude of the contribution of the feature to the model output. The Shapley value of each feature is calculated by the following formula: (14) In the formula: For the first The Shapley value of each feature; M is the number of features, and ! denotes factorial; For the input feature set, ; For features not included Input feature set; For input The model used for forecasting is the LightGBM regression prediction model mentioned above; For features not included A subset of features; For set The number of features; Based on formula (14), the Shapley value of each predictor variable in each sample is calculated, and the mean absolute Shapley value of each predictor variable is calculated. The larger the mean absolute Shapley value, the more significant the contribution of the predictor variable to the runoff prediction result. All predictor variables are sorted in ascending order according to their contribution. The predictor variable with the smallest contribution is eliminated from the current predictor variable set and then re-input into the LightGBM regression model. Time series cross-validation, full sample training and Shapley interpretation are performed. Only one predictor variable is deleted in each iteration. In the recursive elimination process, the model performance evaluation results corresponding to each iteration are recorded. The largest coefficient of determination in the validation set is used as the criterion for the optimal model. The root mean square error (RMSE) and mean absolute error (MAE) are used as auxiliary evaluation indicators to determine the set of predictor variables corresponding to the optimal model. This set is used as the key driving factor set for medium and long-term runoff prediction.

8. The hierarchical screening and recursive optimization method for key driving factors in medium- and long-term runoff forecasting according to claim 7, characterized in that, The MAE calculation formula is as follows: (13)。 9. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes a computer program, it implements the method described in any one of claims 1-8.