Medium and long term runoff prediction method fusing data enhancement technology and machine learning model

By combining data augmentation and machine learning models, the problems of data scarcity and model adaptability in medium- and long-term runoff forecasting have been solved, achieving high-precision and stable runoff forecasting that is suitable for water resource management under complex climatic conditions.

CN121031827APending Publication Date: 2025-11-28CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511181903.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing medium- and long-term runoff forecasting methods face significant challenges in terms of climate change, data scarcity, and model adaptability, resulting in low forecast accuracy, poor stability, and difficulty in adapting to complex climate conditions and data-scarce small and medium-sized watersheds.

Method used

By combining data augmentation techniques with machine learning models, a support vector machine regression model is constructed through factor selection, data augmentation, and model optimization. The SMOTE algorithm is used to generate synthetic samples, and hyperparameters are optimized by combining time lag factors and grid search to improve prediction accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and adaptability of medium- and long-term runoff forecasting, enabling accurate prediction of runoff change trends under complex climatic conditions, providing reliable support for water resource management, and is suitable for runoff forecasting needs in different watersheds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121031827A_ABST
    Figure CN121031827A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of hydrology and water resource and climate prediction crossing, and discloses a medium-and-long-term runoff prediction method fusing a data enhancement technology and a machine learning model, which comprises the following steps: collecting runoff data and climate system index data of a target area, performing factor screening by adopting a replacement accuracy method, the method comprises the following steps: dividing data into a training set and a test set, constructing a time lag factor, carrying out SMOTE data enhancement on the training set, carrying out runoff prediction model construction by using a support vector machine regression model, carrying out parameter optimization by using grid search, calculating an evaluation index to carry out model performance evaluation, and finally predicting and outputting a medium and long-term runoff time sequence according to the model. The method effectively solves the problem of low prediction precision caused by data scarcity and insufficient factor selection in a traditional method. The method overcomes the problems of data scarcity, imbalance and insufficient model generalization ability in medium and long term runoff prediction, and has wide application prospects in the fields of medium and long term runoff prediction and water resource management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of hydrology, water resources and climate prediction, and specifically relates to a medium- and long-term runoff prediction method that integrates data augmentation technology and machine learning models. Background Technology

[0002] Medium- and long-term runoff forecasts are the core basis for water resources planning, water conservancy project scheduling, and flood and drought control decisions. Their accuracy directly affects the efficient utilization and safe management of water resources in a river basin. However, due to the complexity of climate change, the scarcity of hydrological data, and the limitations of forecasting models, existing methods still face significant technical challenges.

[0003] Current medium- and long-term runoff forecasting methods mainly fall into two categories: one is traditional statistical methods based on statistical laws, including time series models such as regression analysis and ARIMA. These methods rely on the assumption of stationarity in historical data, making it difficult to capture the nonlinear and non-stationary runoff characteristics caused by climate change, and their prediction accuracy drops sharply when the data sample size is insufficient. The other category is physical mechanism models, such as rainfall-runoff models like SWAT and HBV, which achieve prediction by simulating the hydrological cycle process. However, they require massive amounts of detailed data on topography, soil, and meteorology, have complex parameter calibration, and high computational costs, limiting their application in small and medium-sized watersheds where data is scarce.

[0004] In recent years, machine learning technology has been introduced into the field of runoff prediction due to its nonlinear fitting capabilities. However, three major bottlenecks still exist in its practical application: First, the blind selection of factors. The correlation between climate system indices (such as the 100 climate indices) and runoff has not been systematically quantified, which easily introduces redundant or irrelevant factors, leading to a decline in the model's generalization ability. Second, the problem of data scarcity and imbalance. Medium- and long-term runoff prediction relies on long-series monthly data, but most watersheds suffer from a shortage of observational data and uneven distribution of samples during wet and dry seasons, which directly affects the training effect of machine learning models. Third, insufficient model adaptability. Existing models are mostly optimized for specific watersheds and have not formed a universal parameter tuning and feature construction scheme, resulting in unstable prediction performance in cross-regional or climate change scenarios.

[0005] To address the aforementioned issues, there is an urgent need for an integrated approach that combines factor screening, data augmentation, and model optimization to improve the accuracy, stability, and practical applicability of medium- and long-term runoff forecasts, and to provide reliable technical support for water resource management under complex climatic conditions. Summary of the Invention

[0006] To overcome the problems of existing technologies, this invention proposes a medium- to long-term runoff prediction method that integrates data augmentation techniques and machine learning models. This method aims to improve the accuracy, robustness, and adaptability of medium- to long-term runoff prediction by combining factor selection, data augmentation, and model optimization, thereby solving prediction challenges caused by data scarcity, insufficient factor selection, and inadequate model generalization ability.

[0007] The objective of this invention is achieved as follows:

[0008] This invention provides a medium- to long-term runoff prediction method that integrates data augmentation techniques and machine learning models, comprising the following steps:

[0009] Step 1: Collect runoff data and climate system index data for the target area.

[0010] Collect runoff data and climate system index data for the target area; the climate system index data is preferably a set of 100 climate system indices published by the China Meteorological Administration.

[0011] The climate system index data and runoff data were aligned by time, missing values ​​and outliers were removed, and the data were standardized.

[0012] Step 2: Factor screening is performed using the permutation accuracy method.

[0013] Based on the permutation accuracy method, the contribution of each climate system index data in step 1 to the runoff in the target area is calculated. The top 20 factors are selected according to their contribution and used as input features for the prediction model.

[0014] Step 3: Divide the data into training and testing sets and construct the time lag factor.

[0015] The selected factors and measured runoff values ​​from step 2 are divided into data sets. 80% are randomly sampled as the training set and the remaining 20% ​​are used as the test set. At the same time, a time lag factor is constructed to generate a composite dataset with stronger predictive power for runoff.

[0016] Step 4: Perform SMOTE data augmentation on the training set.

[0017] The SMOTE algorithm is used to augment the selected time lag factors and measured runoff values ​​in the training set divided in step 3, resulting in an augmented composite dataset.

[0018] Step 5: Construct a runoff prediction model using a support vector machine model.

[0019] Based on the enhanced composite dataset generated in step 4, a support vector machine regression model is constructed as a runoff prediction tool.

[0020] Step 6: Optimize parameters using grid search

[0021] The performance of the support vector machine regression model constructed in step 5 was evaluated using the 5-fold cross-validation method, and the hyperparameters were adjusted by the grid search method to find the optimal parameter combination.

[0022] Step 7: Calculate evaluation metrics and conduct model performance evaluation.

[0023] Input the test set data into the support vector machine regression model trained in step 6, calculate the model's prediction metrics, and evaluate the model's overall performance and feasibility for practical application.

[0024] Step 8: Predict medium- and long-term runoff time series based on the model.

[0025] Using the optimized support vector machine regression model in step 6, and based on the top 20 contributing factors selected in step 2, the latest climate system index data for the target region is input, and the predicted monthly runoff values ​​for the future are output to generate a medium- to long-term runoff prediction time series.

[0026] Furthermore, in step 1, the data is standardized using the Z-Score standardization method, which specifically includes: subtracting the mean from each data point and then dividing by the standard deviation, so that the mean of all feature data is 0 and the standard deviation is 1, in order to eliminate the influence of dimensional differences on model training.

[0027] Furthermore, in step 2, the contribution of climate system index data is calculated using the permutation accuracy method, and the stability and rationality of factor selection are verified by multiple random samplings. Specifically, this includes:

[0028] (1) Use the climate system index data obtained in step 1 as the input feature, denoted as matrix X = {x1, x2, ..., x...} n}, where x i This represents the time series data of the i-th climate system index, where n is the total number of features; the measured runoff value is used as the prediction target, denoted as vector Y = {y1, y2, ..., y...} m}, where y j This represents the measured runoff value at the j-th time point, where m is the total number of samples;

[0029] (2) Using matrix X and vector Y as input, construct an initial prediction model (such as a random forest regression model or a support vector machine regression model), and use 5-fold cross-validation to calculate the baseline performance index of the model on the test set, denoted as the baseline error E0 (mean squared error RMSE or coefficient of determination R can be selected). 2 wait);

[0030] (3) For each feature x in the input feature matrix X iThe permutation process is performed sequentially, that is, the time series order of the feature is randomly shuffled to obtain the permuted feature vector x. i The original matrix X is then input into the model using the permuted feature matrix X', while keeping other features unchanged. The prediction error of the model is then recalculated and denoted as E. i ;

[0031] (4) Define feature x i The substitution contribution I i The calculation formula is as follows:

[0032] I i =E i -E0

[0033] Among them, I i Indicates feature x i The increase in error caused by the substitution; a larger value indicates a higher contribution of the feature to the model's predictive performance; E i This indicates that when feature x is used... i After random permutation (shuffling the order), the prediction error is recalculated by the model.

[0034] (5) For each feature x i Repeat the permutation operation in steps (3) and (4) T times (T≥30), and take the average contribution value as the final feature importance score, denoted as:

[0035] S i = (1 / T)×Σ(E) it -E0), t=1,2,...,T

[0036] In the formula, S i The final feature importance score is represented by T, where T represents the number of permutation operations, and E represents the number of permutation operations. it This indicates that in the t-th experiment, for feature x i The prediction error of the prediction model after the permutation;

[0037] (6) According to the feature score S i Sort the features from highest to lowest, and select the top k features (k=20) as the final input feature set, denoted as:

[0038] X = {x1, x2, ..., x} k}

[0039] (7) To verify the stability of the screening results, the Bootstrap random resampling method was used. 80% of the data in the original sample were randomly selected and the above feature contribution calculation process was repeated N times (N≥50). The selection frequency of each feature in all resampling was counted, and features with selection frequency lower than the set threshold (e.g., 70%) were removed to ensure that the selected features are stable and representative.

[0040] Furthermore, in step 3, the time lag factor includes the selected factors and runoff data for the target area from the previous 1 month to the previous 12 months.

[0041] Furthermore, in step 4, the SMOTE algorithm is used for data augmentation. Different augmentation ratios and random seeds are set, and the optimal augmentation parameters are selected through multiple rounds of experiments. Specifically, this includes:

[0042] (1) Based on the distribution of the training set labels, the measured runoff values ​​are divided into low flow, medium flow and high flow categories according to the set thresholds Z1 and Z2, and the differences in the number of samples in each category are analyzed.

[0043] (2) For categories with a small number of samples, oversampling is performed using the SMOTE algorithm: For each sample point x to be augmented i Find the M nearest neighbor samples in the feature space, and randomly select a sample point x from them. j Calculate the difference vector d = x j -x i Generate a new sample x':

[0044] x'=x i +α×d

[0045] In the formula, α∈(0,1) represents the random interpolation coefficients;

[0046] (3) Set different augmentation ratios R∈{0.5,1.0,1.5,2.0}, representing the ratio of the number of new samples generated to the number of original minority class samples; the random seed S∈{1,2,...,10} controls the randomness of interpolation; conduct multiple rounds of experiments on each set of parameters (R,S) to evaluate the predictive performance of the augmented dataset (such as RMSE, NSE), and select the optimal augmentation parameters (R,S)**;

[0047] (4) The training set is augmented using the optimal augmentation parameters (R,S) to obtain the augmented training feature matrix and the augmented training label vector, which are used as the dataset for model training.

[0048] (5) The test set data does not participate in data augmentation, maintains the original data distribution, and is used for independent verification of model performance to avoid data leakage.

[0049] Furthermore, in steps 5 and 6, the support vector machine regression model uses the radial basis function (RBF) and searches for the optimal combination within the range of C values ​​(0.1, 1000) and γ values ​​(0.01, 100) using a grid search method.

[0050] The grid search method specifically includes:

[0051] (1) Set the search range of the penalty parameter C and the kernel function parameter γ, C∈{0.1,1,10,100,500,1000}, γ∈{0.01,0.1,1,10,50,100}, to form a parameter grid of C and γ;

[0052] (2) For each set of parameter combinations in the grid (C) i ,γ j Train a support vector machine regression model, and use the training set data to calculate the model performance metrics for this combination, including mean squared error (RMSE) and Nash efficiency coefficient (NSE), denoted as RMSE(C). i ,γ j ) and NSE(C i ,γ j );

[0053] (3) Five-fold cross-validation was used to validate each parameter combination (C) i ,γ j To verify the results, the training set data is divided into 5 subsets. One subset is used as the validation set and the other 4 subsets are used as the training set. After 5 iterations, the average performance index is taken as the evaluation value of the parameter combination.

[0054] (4) Among all parameter combinations, select the parameter combination with the smallest RMSE and the largest NSE under cross-validation, denoted as (C*, γ*), as the optimal parameters for the final model training;

[0055] (5) Apply the optimal parameters (C*, γ*) to the support vector machine regression model, retrain the full training data, and obtain the final prediction model.

[0056] This invention integrates data augmentation techniques, factor selection methods, and machine learning models to construct a highly efficient solution suitable for medium- and long-term runoff forecasting, which has the following significant advantages compared to existing technologies:

[0057] 1. Improved Prediction Accuracy and Enhanced Model Robustness: Addressing the issue of low prediction accuracy in scenarios with scarce data and imbalanced samples using traditional methods, this invention generates synthetic samples through the SMOTE algorithm. This effectively expands the distribution range of small sample data while maintaining consistency between the synthetic and original data distributions, reducing the risk of overfitting due to insufficient data. Combined with 20 high-contribution climate factors selected using the permutation accuracy method, it ensures a strong correlation between input features and runoff, providing a high-quality input foundation for the model.

[0058] The support vector machine (SVM) regression model, combined with 5-fold cross-validation and grid search optimization of hyperparameters, further enhances the ability to capture nonlinear runoff variation patterns. Experimental validation shows that, through RMSE and NSE index evaluation, the model prediction error is significantly reduced, and it has stronger resistance to runoff fluctuations under complex climatic conditions, demonstrating significantly enhanced robustness.

[0059] 2. Optimized feature construction and enhanced predictive capabilities: An innovative approach introduces time lag factors (screening factors and runoff data from the previous 1 to 12 months) to construct a composite feature dataset. This fully explores the temporal correlation of runoff changes and solves the problem of traditional models neglecting the lag effect of hydrological processes. This feature engineering method enables the model to better capture seasonal and periodic runoff patterns, improving the accuracy of medium- and long-term (e.g., monthly) predictions.

[0060] The substitution accuracy factor selection method avoids the blind selection of factors by quantifying the contribution of climate indices to runoff, ensuring the representativeness and stability of the selected features, reducing the interference of redundant information on the model, and improving prediction efficiency.

[0061] 3. Enhanced applicability and expanded application scenarios: This invention offers greater flexibility in data quality requirements. Through data augmentation techniques and standardized processing, it can adapt to the runoff prediction needs of different watersheds (especially small and medium-sized watersheds with scarce data), solving the problem of physical models' dependence on massive amounts of detailed data. Simultaneously, the model parameters are dynamically optimized through grid search, adapting to the climatic characteristics of different regions and exhibiting strong cross-regional applicability.

[0062] The forecast results are output in the form of medium- to long-term monthly runoff time series, which can directly provide quantitative basis for water resource planning, water conservancy project scheduling, and flood control and drought relief decision-making. For example, the runoff forecast for the next 12 months can support the formulation of reservoir water storage and scheduling plans, and the runoff trend analysis under the influence of climate change can assist in the sustainable management of watershed water resources, demonstrating significant application value.

[0063] 4. Efficient technical process, easy to implement in practice: The entire methodology (data preprocessing → factor selection → data augmentation → model training → prediction output) forms a standardized framework. Combined with the scikit-learn library, it enables efficient deployment of SVM models and grid search, with low computational cost and convenient operation, making it easy to promote in practical applications in hydrological business departments. The folded cross-validation and dynamic hyperparameter adjustment mechanism ensure the objectivity of model performance evaluation and the optimality of parameter selection, reducing the influence of human experience on prediction results and making the method more scientific and reproducible. Attached Figure Description

[0064] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0065] Figure 1 This is a flowchart of the method described in the embodiments of the present invention;

[0066] Figure 2 This is a time series comparison chart of monthly runoff rates.

[0067] Figure 3 This is a scatter plot comparing measured and simulated monthly runoff rates.

[0068] Figure 4 This is a comparison chart of monthly runoff time series during the verification period;

[0069] Figure 5 This is a scatter plot comparing measured and simulated monthly runoff during the verification period. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0071] Example 1:

[0072] like Figure 1 As shown, this embodiment provides a method for predicting medium- and long-term runoff in a target area based on monthly measured runoff data and a set of climate system indices. It offers a method that integrates data augmentation techniques and machine learning models for medium- and long-term runoff prediction, including the following steps:

[0073] Step 1: Collect runoff data and climate system index data for the target area.

[0074] Based on the hydrological station observation records of the target area, monthly measured runoff data were collected, and a set of 100 climate system indices released by the China Meteorological Administration were obtained, including factors such as NINO3.4, PDO, and AO.

[0075] The data were initially screened to remove missing and outlier values. If a factor was missing or had more than 30% of its values ​​being outliers, it was disregarded. After aligning the runoff data with the climate system index data by time, the data were normalized using the Z-Score standardization method. Each data point was subtracted from its mean and then divided by its standard deviation to ensure that the mean of all factor data was 0 and the standard deviation was 1, thus eliminating the influence of different units of measurement on model training.

[0076] Step 2: Factor screening is performed using the permutation accuracy method.

[0077] The permutation accuracy method is used to evaluate the contribution of each climate system index to runoff. Specifically, a machine learning model, typically a random forest or support vector machine model, is first trained by inputting all factors and runoff labels, and its baseline accuracy (or other performance metrics, such as mean squared error) on the test set is calculated. Then, for each feature (factor), its order is randomly shuffled to ensure that the relationship between the feature and its target variable is disrupted. The results are then re-predicted on the shuffled dataset, and the new accuracy of the model is calculated. The difference in accuracy before and after the permutation (e.g., a decrease in accuracy or an increase in error) represents the importance of the feature; the greater the loss, the greater the contribution of the feature to the model's predictive performance. Through multiple randomized experiments, the top 20 climate factors (such as NINO3.4, AO, SOI, etc.) contributing to runoff prediction in the target area are selected as the input feature set for the model.

[0078] Specifically, it includes:

[0079] (1) Use the climate system index data obtained in step 1 as the input feature, denoted as matrix X = {x1, x2, ..., x...} n}, where x i This represents the time series data of the i-th climate system index, where n is the total number of features; the measured runoff value is used as the prediction target, denoted as vector Y = {y1, y2, ..., y...} m}, where y j This represents the measured runoff value at the j-th time point, where m is the total number of samples;

[0080] (2) Using matrix X and vector Y as input, construct an initial prediction model (such as a random forest regression model or a support vector machine regression model), and use 5-fold cross-validation to calculate the baseline performance index of the model on the test set, denoted as the baseline error E0 (mean squared error RMSE or coefficient of determination R can be selected). 2 wait);

[0081] (3) For each feature x in the input feature matrix X i The permutation process is performed sequentially, that is, the time series order of the feature is randomly shuffled to obtain the permuted feature vector x. i The original matrix X is then input into the model using the permuted feature matrix X', while keeping other features unchanged. The prediction error of the model is then recalculated and denoted as E. i ;

[0082] (4) Define feature x i The substitution contribution I i The calculation formula is as follows:

[0083] I i =E i -E0

[0084] Among them, I i Indicates feature x i The increase in error caused by the substitution; a larger value indicates a higher contribution of the feature to the model's predictive performance; E i This indicates that when feature x is used... i After random permutation (shuffling the order), the prediction error is recalculated by the model.

[0085] (5) For each feature x i Repeat the permutation operation in steps (3) and (4) T times (T≥30), and take the average contribution value as the final feature importance score, denoted as:

[0086] S i = (1 / T)×Σ(E) it -E0), t=1,2,...,T

[0087] In the formula, S i The final feature importance score is represented by T, where T represents the number of permutation operations, and E represents the number of permutation operations. it This indicates that in the t-th experiment, for feature x i The prediction error of the prediction model after the permutation;

[0088] (6) According to the feature score S i Sort the features from highest to lowest, and select the top k features (k=20) as the final input feature set, denoted as:

[0089] X = {x1, x2, ..., x} k}

[0090] (7) To verify the stability of the screening results, the Bootstrap random resampling method was used. 80% of the data in the original sample were randomly selected and the above feature contribution calculation process was repeated N times (N≥50). The selection frequency of each feature in all resampling was counted, and features with selection frequency lower than the set threshold (e.g., 70%) were removed to ensure that the selected features are stable and representative.

[0091] Step 3: Divide the data into training and testing sets and construct the time lag factor.

[0092] The selected climate factors and measured runoff values ​​from step 2 are divided into data sets. 80% are randomly sampled as the training set, and the remaining 20% ​​is used as the test set. At the same time, a time lag factor is constructed, taking the climate factor values ​​from the previous 1 month to the previous 12 months (i.e., the 20 climate factors selected in step 2) and the corresponding runoff data as features to generate a composite feature dataset, further improving the ability to capture runoff variation patterns.

[0093] Step 4: Perform SMOTE data augmentation on the training set.

[0094] Data augmentation is performed on the selected time lag factors and measured runoff values ​​in the training set divided in step 3 to obtain an augmented composite dataset.

[0095] Data augmentation uses the SMOTE (Synthetic Minority Over-sampling Technique) algorithm to balance factor data and runoff label datasets, especially when faced with scarce data and imbalanced samples. SMOTE augments minority class samples by generating new synthetic samples.

[0096] Specifically, for each minority class sample,

[0097] (1) Based on the distribution of the training set labels, the measured runoff values ​​are divided into low flow, medium flow and high flow categories according to the set thresholds Z1 and Z2, and the differences in the number of samples in each category are analyzed.

[0098] (2) For categories with a small number of samples, oversampling is performed using the SMOTE algorithm: For each sample point x to be augmented i Find the M nearest neighbor samples in the feature space, and randomly select a sample point x from them. j Calculate the difference vector d = x j -x i Generate a new sample x':

[0099] x'=x i +α×d

[0100] In the formula, α∈(0,1) represents the random interpolation coefficients;

[0101] (3) Set different augmentation ratios R∈{0.5,1.0,1.5,2.0}, representing the ratio of the number of new samples generated to the number of original minority class samples; the random seed S∈{1,2,...,10} controls the randomness of interpolation; conduct multiple rounds of experiments on each set of parameters (R,S) to evaluate the predictive performance of the augmented dataset (such as RMSE, NSE), and select the optimal augmentation parameters (R,S)**;

[0102] (4) The training set is augmented using the optimal augmentation parameters (R,S) to obtain the augmented training feature matrix and the augmented training label vector, which are used as the dataset for model training.

[0103] (5) The test set data does not participate in data augmentation, maintains the original data distribution, and is used for independent verification of model performance to avoid data leakage.

[0104] Step 5: Construct a runoff prediction model using a support vector machine regression model.

[0105] Based on the augmented composite dataset generated in step 4, a Support Vector Machine (SVM) regression model is trained. SVM regression measures the deviation between predicted and true values ​​by defining an ∈-insensitive band (epsilon-tube). If a data point falls within this insensitive band, no loss is calculated; if it exceeds this band, the loss is calculated based on the deviation. Furthermore, it achieves a balance by optimizing the following two objectives: ① Minimizing model complexity: ensuring the prediction function is as simple as possible. ② Controlling the error range: tolerating error within the ∈-insensitive band and optimizing deviation outside the range.

[0106] Step 6: Optimize parameters using grid search

[0107] The model performance was evaluated using a 5-fold cross-validation method. The support vector machine (SVM) regression model used a radial basis function (RBF) kernel function. The optimal combination of C values ​​(0.1, 1000) and γ values ​​(0.01, 100) was found through a grid search method to optimize the data distribution and model parameters.

[0108] The core idea of ​​grid search is to traverse all possible combinations of hyperparameters, train the model, and evaluate its performance. Specific steps include:

[0109] (1) Set the search range of the penalty parameter C and the kernel function parameter γ, C∈{0.1,1,10,100,500,1000}, γ∈{0.01,0.1,1,10,50,100}, to form a parameter grid of C and γ;

[0110] (2) For each set of parameter combinations in the grid (C) i ,γ j Train a support vector machine regression model, and use the training set data to calculate the model performance metrics for this combination, including mean squared error (RMSE) and Nash efficiency coefficient (NSE), denoted as RMSE(C). i ,γ j ) and NSE(C i ,γ j );

[0111] (3) Five-fold cross-validation was used to validate each parameter combination (C) i ,γ j To verify the results, the training set data is divided into 5 subsets. One subset is used as the validation set and the other 4 subsets are used as the training set. After 5 iterations, the average performance index is taken as the evaluation value of the parameter combination.

[0112] (4) Among all parameter combinations, select the parameter combination with the smallest RMSE and the largest NSE under cross-validation, denoted as (C*, γ*), as the optimal parameters for the final model training;

[0113] (5) Apply the optimal parameters (C*, γ*) to the support vector machine regression model, retrain the full training data, and obtain the final prediction model.

[0114] Step 7: Calculate evaluation metrics to assess model performance.

[0115] Input the test set data into the model trained in step 6, and calculate the model's prediction metrics on the test set, including RMSE and NSE. RMSE measures the overall prediction error of the model, while NSE evaluates the model's fit to runoff change trends, ensuring the model's feasibility in practical applications.

[0116] Step 8: Output medium- and long-term runoff time series based on model predictions.

[0117] Using the optimized SVM model and the latest climate system index data from 2023-2024 as input, the model outputs predicted future runoff values ​​for the target area, generating a medium- to long-term runoff forecast time series. To ensure the real-time nature and accuracy of the prediction results, a sliding window method is used to dynamically update the model, adapting it to the long-term changing trends of climate factors, thus providing reliable predictive support for watershed water resources management.

[0118] In summary, this invention, based on data augmentation techniques, support vector machine regression models, and factor selection methods, constructs an efficient method suitable for medium- and long-term runoff prediction, solving the problems of data scarcity, sample imbalance, and insufficient factor selection. By introducing the SMOTE algorithm to augment minority class samples and combining it with time lag factors, the model's ability to capture dynamic runoff patterns is improved. In the context of climate change, this method can accurately predict future runoff trends, providing strong support for water resource management and planning, and has broad application prospects.

[0119] Application examples:

[0120] Taking the Danjiangkou watershed as an application example, and combining actual observation data, the medium- and long-term runoff prediction method based on the fusion of data augmentation technology and machine learning model described in Embodiment 1 of this invention is used to simulate and predict monthly runoff. The specific steps are as follows:

[0121] First, measured runoff data from 2009 to 2017 and corresponding data from 100 climate system indices were selected as calibration period data, and data from 2018 to 2020 were selected as validation period data. Following the method steps of this invention, data preprocessing and feature selection were completed, and the top 20 climate system indices contributing most to runoff prediction were selected as input features.

[0122] During the data augmentation phase, the SMOTE algorithm was used to augment the high-water season data, which had a smaller sample size in the periodic training set. This balanced the sample size across different flow intervals and enhanced the model's ability to fit extreme high-flow conditions. By setting different augmentation ratios and random seeds, and combining the results of cross-validation, the final combination of augmentation parameters was determined (augmentation ratio R = 1.0, random seed S = 5).

[0123] Based on the enhanced training data, a support vector machine regression model was constructed. A radial basis function (RBF) kernel was used, and the parameters were optimized within the range of C (0.11000) and γ (0.01100) using a grid search method. Finally, C = 100 and γ = 0.1 were selected as the optimal parameter combination. The predictive performance of the model on the training data was evaluated using 5-fold cross-validation.

[0124] The optimized SVM model was applied to predict the validation period data, and the results are as follows: Figure 2-5 As shown.

[0125] The results show that during the calibration period (2009-2017), the model predictions and observed runoff trends are in good agreement, with a Nash efficiency coefficient (NSE) of 0.958 and a root mean square error (RMSE) of 14.032 cubic meters per second (e.g., Figure 2 , 3 As shown). During the validation period (2018-2020), the model prediction results also showed good fit, with an NSE of 0.843 and an RMSE of 34.958 cubic meters per second (as shown). Figure 4 , 5 (As shown).

[0126] The medium- to long-term runoff prediction method based on the fusion of data augmentation technology and machine learning models described in this invention significantly improves the accuracy and robustness of runoff prediction under conditions of data scarcity and sample imbalance. It demonstrates excellent adaptability and practical application value, particularly in watershed environments characterized by alternating wet and dry seasons and frequent extreme events. This example verifies the broad applicability of the method under complex climatic conditions and varying data scales, providing reliable medium- to long-term runoff prediction support for watershed water resources planning and management.

[0127] Finally, it should be noted that the above is only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention (such as the application of various formulas, the order of steps, etc.) without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A medium- to long-term runoff prediction method integrating data augmentation technology and machine learning models, characterized in that, The method includes the following steps: Step 1: Collect runoff data and climate system index data for the target area. Collect runoff data and climate system index data for the target area; the climate system index data is preferably a set of 100 climate system indices published by the China Meteorological Administration. The climate system index data and runoff data were aligned by time, missing values ​​and outliers were removed, and the data were standardized. Step 2: Factor screening is performed using the permutation accuracy method. Based on the permutation accuracy method, the contribution of each climate system index data in step 1 to the runoff in the target area is calculated, and the top 20 factors are selected according to the contribution ranking as input features of the prediction model. Step 3: Divide the data into training and testing sets and construct the time lag factor. The selected factors and measured runoff values ​​in step 2 are divided into data sets. 80% are randomly sampled as the training set and the remaining 20% ​​are used as the test set. At the same time, a time lag factor is constructed to generate a composite dataset with stronger predictive power for runoff. Step 4: Perform SMOTE data augmentation on the training set. The SMOTE algorithm was used to augment the selected time lag factors and measured runoff values ​​in the training set partitioned in step 3 to obtain an augmented composite dataset. Step 5: Construct a runoff prediction model using a support vector machine model. Based on the enhanced composite dataset generated in step 4, a support vector machine regression model is constructed as a runoff prediction tool. Step 6: Optimize parameters using grid search The performance of the support vector machine regression model constructed in step 5 was evaluated using the 5-fold cross-validation method, and the hyperparameters were adjusted by the grid search method to find the optimal parameter combination. Step 7: Calculate evaluation metrics and conduct model performance evaluation. Input the test set data into the support vector machine regression model trained in step 6, calculate the model's prediction metrics, and evaluate the model's overall performance and feasibility for practical application. Step 8: Predict medium- and long-term runoff time series based on the model. Using the optimized support vector machine regression model in step 6, and based on the top 20 contributing factors selected in step 2, the latest climate system index data for the target region is input, and the predicted monthly runoff values ​​for the future are output to generate a medium- to long-term runoff prediction time series.

2. The method according to claim 1, characterized in that, In step 1, the data is standardized using the Z-Score standardization method, which specifically includes: Subtract the mean from each data point and divide by the standard deviation to make the mean of all feature data 0 and the standard deviation 1, thus eliminating the impact of dimensional differences on model training.

3. The method according to claim 1, characterized in that, In step 2, the contribution of climate system index data is calculated using the permutation accuracy method, and the stability and rationality of factor selection are verified by multiple random samplings. Specifically, this includes: (1) Use the climate system index data obtained in step 1 as the input feature, denoted as matrix X = {x1, x2, ..., x...} n }, where x i This represents the time series data of the i-th climate system index, where n is the total number of features; the measured runoff value is used as the prediction target, denoted as vector Y = {y1, y2, ..., y...} m }, where y j This represents the measured runoff value at the j-th time point, where m is the total number of samples; (2) Using matrix X and vector Y as input, construct an initial prediction model, and use 5-fold cross-validation to calculate the benchmark performance index of the model on the test set, denoted as the benchmark error E0; (3) For each feature x in the input feature matrix X i The permutation process is performed sequentially, that is, the time series order of the feature is randomly shuffled to obtain the permuted feature vector x. i The original matrix X is then input into the model using the permuted feature matrix X', while keeping other features unchanged. The prediction error of the model is then recalculated and denoted as E. i ; (4) Define feature x i The substitution contribution I i The calculation formula is as follows: I i =E i -E0 In the formula, I i Indicates feature x i The increase in error caused by the substitution; E i This indicates that when feature x is used... i After random permutation, the prediction error is recalculated by re-entering the model. (5) For each feature x i Repeat the permutation operation in steps (3) and (4) T times, and take the average contribution value as the final feature importance score, denoted as: S i =(1 / T)×Σ(E it -E0),t=1,2,...,T In the formula, S i The final feature importance score is represented by T, where T represents the number of permutation operations, and E represents the number of permutation operations. it This indicates that in the t-th experiment, for feature x i The prediction error of the prediction model after the permutation; (6) According to the feature score S i Sort the features from highest to lowest, and select the top k features as the final input feature set, denoted as: X={x1,x2,…,x k } (7) To verify the stability of the screening results, the Bootstrap random resampling method was used. 80% of the data in the original sample were randomly selected and the above feature contribution calculation process was repeated N times. The selection frequency of each feature in all resampling was counted, and features with selection frequency lower than the set threshold were removed to ensure that the selected features are stable and representative.

4. The method according to claim 1, characterized in that, In step 3, the time lag factor includes the selected factors and runoff data for the target area from the previous 1 month to the previous 12 months.

5. The method according to claim 1, characterized in that, In step 4, the SMOTE algorithm is used for data augmentation. Different augmentation ratios and random seeds are set, and the optimal augmentation parameters are selected through multiple rounds of experiments. Specifically, this includes: (1) Based on the distribution of the training set labels, the measured runoff values ​​are divided into low flow, medium flow and high flow categories according to the set thresholds Z1 and Z2, and the differences in the number of samples in each category are analyzed. (2) Oversampling using the SMOTE algorithm: For each sample point x to be augmented... i Find the M nearest neighbor samples in the feature space, and randomly select a sample point x from them. j Calculate the difference vector d = x j -x i Generate a new sample x': x'=x i +α×d In the formula, α is the random interpolation coefficient; (3) Set different enhancement ratios R∈{0.5,1.0,1.5,2.0}, representing the ratio of the number of new samples generated to the number of original minority class samples; random seed S∈{1,2,...,10} controls the randomness of interpolation, conduct multiple rounds of experiments on each set of parameters (R,S) to evaluate the predictive performance of the enhanced dataset, and select the optimal enhancement parameters (R,S)**; (4) The training set is augmented using the optimal augmentation parameters (R,S) to obtain the augmented training feature matrix and the augmented training label vector, which are used as the dataset for model training. (5) The test set data does not participate in data augmentation, maintains the original data distribution, and is used for independent verification of model performance to avoid data leakage.

6. The method according to claim 1, characterized in that, In steps 5 and 6, the support vector machine regression model uses the radial basis function (RBF) and searches for the optimal combination of C values ​​(0.1, 1000) and γ values ​​(0.01, 100) using a grid search method.

7. The method according to claim 6, characterized in that, The grid search method specifically includes: (1) Set the search range of the penalty parameter C and the kernel function parameter γ, C∈{0.1,1,10,100,500,1000}, γ∈{0.01,0.1,1,10,50,100}, to form a parameter grid of C and γ; (2) For each set of parameter combinations in the grid (C) i ,γ j Train a support vector machine regression model, and use the training set data to calculate the model performance metrics for this combination, including mean squared error (RMSE) and Nash efficiency coefficient (NSE), denoted as RMSE(C). i ,γ j ) and NSE(C i ,γ j ); (3) Five-fold cross-validation was used for each parameter combination (C i ,γ j To verify the results, the training set data is divided into 5 subsets. One subset is used as the validation set and the other 4 subsets are used as the training set. After 5 iterations, the average performance index is taken as the evaluation value of the parameter combination. (4) Among all parameter combinations, select the parameter combination with the smallest RMSE and the largest NSE under cross-validation, denoted as (C*, γ*), as the optimal parameters for the final model training; (5) Apply the optimal parameters (C*, γ*) to the support vector machine regression model, retrain the full training data, and obtain the final prediction model.