Hybrid machine learning-based method for predicting hydropower under climate change

By establishing a hydrology-reservoir-hydropower coupling model and machine learning algorithms, the uncertainties and computational complexity of hydropower forecasting under climate change in existing technologies have been solved. This has enabled refined simulation and rapid prediction of hydropower output under climate change, reduced computational complexity and uncertainty, and improved the practicality of the assessment method.

CN122334877APending Publication Date: 2026-07-03CHINA THREE GORGES PROJECTS DEV CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA THREE GORGES PROJECTS DEV CO LTD
Filing Date
2026-05-20
Publication Date
2026-07-03

Smart Images

  • Figure CN122334877A_ABST
    Figure CN122334877A_ABST
Patent Text Reader

Abstract

This invention discloses a method for hydropower prediction under climate change that integrates machine learning, belonging to the field of hydropower data management technology. It includes: simulating the hydropower response under future climate change using a finite number of representative samples in a target region through a physical process-based hydrological-reservoir-hydropower coupling model; constructing a training dataset containing complex nonlinear response characteristics while ensuring the reliability of the model's physical principles and results; then, using machine learning algorithms to establish an efficient generalization model between key climate change features and hydropower generation; finally, applying this model to a larger sample of future climate prediction data covering a wider range of climate change scenarios to achieve rapid and robust hydropower output prediction and uncertainty quantification. This provides more robust and reliable technical support for hydropower prediction under climate change scenarios, as well as for uncertainty quantification and risk decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydropower data management technology, and in particular to a method for hydropower prediction under climate change that integrates machine learning. Background Technology

[0002] Since the Industrial Revolution, the continuous increase in atmospheric greenhouse gas concentrations has led to significant global warming, which has significantly intensified the Earth's water cycle and is altering the spatiotemporal distribution of global surface water resources. This change has profoundly impacted reservoir operation and the efficiency and stability of hydropower generation. Current methods for quantifying hydropower output under climate change primarily employ a top-down assessment framework, based on global climate models coupled with hydrological models and reservoir operation and power generation models, to assess the impact of future climate change on hydrology, water resources, and hydropower generation. This provides a possibility for quantitatively assessing the potential impacts of climate change on a centennial scale, and has significant guiding significance for the operation of major water conservancy projects and energy development planning.

[0003] Recent studies have shown that hydropower output exhibits a complex nonlinear response to surface runoff and climate change. Accurately quantifying the potential impact of climate change on hydropower production requires a detailed characterization of surface hydrophysical processes and reservoir operation, necessitating comprehensive data support and the integration of multidisciplinary knowledge. Meanwhile, rapidly developing machine learning methods possess powerful capabilities for processing high-dimensional, complex data, automatically uncovering patterns, and continuously adapting, offering a new approach to addressing these challenges. Machine learning methods can efficiently integrate multi-source data from meteorology, hydrology, and hydropower, significantly improving simulation speed and prediction accuracy, thus providing an efficient and accurate data-driven solution for intelligent prediction of hydropower systems. With the accumulation of observational data and high-resolution climate forecasting data, machine learning is expected to achieve higher-precision output predictions in hydropower simulations and support large-scale scenario simulations, providing interpretable and scalable analytical tools for assessing climate change-related uncertainties. This will further promote the deep integration of hydropower scientific research and intelligent operation management.

[0004] For example, the Chinese invention patent with publication number CN120562799A discloses a multi-objective optimization hydropower ecological scheduling decision system and method. The system includes modules for data acquisition, ecological model, intelligent decision engine, scheduling execution, effect evaluation, and knowledge base. The data acquisition module collects multi-source data, the ecological model module generates training datasets, the intelligent decision engine performs ecological process modeling and probability prediction based on deep learning architecture and spatiotemporal feature extraction, generates scheduling decisions, the scheduling execution module controls the operation of hydropower projects, the effect evaluation module monitors ecological and economic effects, and the knowledge base module stores historical experience and provides optimization suggestions. The system improves the accuracy of ecosystem simulation, especially during sudden changes in flow.

[0005] The above-mentioned technology has at least the following technical problems: Existing assessment processes based on global climate models coupled with hydrological models and reservoir operation and power generation models involve lengthy chains of processes. They also suffer from uncertainties in scenario assumptions, climate models, and simulations of water resources and hydropower production systems. Furthermore, they face challenges such as high requirements for basic data, complex modeling processes, and large computational demands, making them difficult to apply to comprehensive climate change scenarios and complex watershed and hydraulic engineering facilities. When the number of climate change prediction data models is limited, the conclusions drawn are highly uncertain. Existing machine learning-based reservoir operation models have yet to achieve effective integration with hydropower output predictions under climate change. Summary of the Invention

[0006] This invention provides a method for predicting hydropower under climate change by incorporating machine learning, the technical solution of which is as follows: The parameters of the hydrological model were calibrated and validated, and historical runoff simulation data from representative hydrological stations within the target area were obtained. The hydrological model was coupled with reservoir scheduling and hydropower models. Historical runoff simulation data from representative hydrological stations within the target area were used to drive the reservoir scheduling model, and parameter optimization was performed on the reservoir scheduling and hydropower models to establish a hydrological-reservoir-hydropower coupled model. The first sample of climate change prediction datasets for the target area was obtained and analyzed, and the future climate change prediction datasets of the first sample of the target area were downscaled and biased. The climate change data of the first sample under different future climate scenarios were then analyzed. The project uses a hydrological-reservoir-hydropower coupled model driven by projected data to obtain the power generation time series of the first sample in the target area under future climate scenarios. Based on the power generation time series of the first sample in the target area under future climate scenarios, combined with hydropower generation and climate change characteristics, a generalized machine learning model for hydropower prediction is established using machine learning algorithms. The project then obtains the future climate change prediction data of the second sample in the target area, drives the generalized machine learning model for hydropower prediction based on climate change characteristics, and obtains the future hydropower prediction results under the second sample in the target area, quantifying the uncertainty of the future hydropower prediction results under the second sample in the target area.

[0007] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention provides a method for hydropower prediction under climate change that integrates machine learning. First, for a representative first sample in the target region, a refined simulation of the hydropower response under future climate change is performed using a hydrological-reservoir-hydropower coupling model based on physical processes. While ensuring the reliability of the model's physical principles and results, a training dataset containing complex nonlinear response characteristics is constructed. Next, a highly efficient generalization model between key climate change features and hydropower generation is established using machine learning algorithms. Finally, this model is applied to a second sample of future climate prediction data covering a larger scale and broader range of climate change scenarios, achieving rapid and robust hydropower output prediction and uncertainty quantification. This technical approach significantly reduces the computational complexity and workload of directly preprocessing large-scale climate scenario data and simulating reservoir scheduling, and reduces reliance on interdisciplinary expertise, thereby improving the practicality and operability of the assessment method. Simultaneously, the constructed machine learning model has good scalability and can be quickly applied to large-scale climate scenario sets and even future new scenarios, reducing the uncertainty brought by a single scenario or model, and providing more robust prediction results for uncertainty and risk decision-making based on integrated climate change data.

[0008] 2. This invention systematically compares and filters different climate models and various downscaling methods in the first sample of climate change prediction datasets for the target region. Using actual climate observation data as a reference, it quantitatively evaluates the simulation results of each combination, thereby comprehensively determining the optimal combination of climate model and downscaling method in terms of both error control and consistency of change trends. By simultaneously examining the root mean square error and correlation coefficient, it avoids the bias that may arise from relying on a single evaluation indicator, making the selected climate scenario closer to the real climate state in terms of mean level, variability characteristics, and interannual fluctuations. This provides more reliable climate input conditions for subsequent hydrological models and reservoir operation models. Therefore, it effectively reduces the uncertainty introduced by climate model selection and downscaling, and improves the transmissibility and stability of climate change signals in hydropower simulation.

[0009] 3. This invention constructs a machine learning training dataset by combining the hydropower generation sequence obtained from the first sample of the target area with characteristic parameters such as the annual average value, interannual coefficient of variation, and seasonal distribution concentration of key climate factors. It uses machine learning methods to establish a generalized machine learning model for hydropower prediction between power generation and the statistical characteristic parameters of key climate factors, thus establishing an effective connection between the physical mechanism model and the data-driven conceptual model. When applied to the hydropower prediction of the second sample, it significantly reduces the input requirements of the prediction model, avoids the complex calculation process of statistical downscaling of climate models and coupling of hydrology-reservoir scheduling-hydropower models, and the decision tree algorithm used can effectively preserve the complex nonlinear response relationship between hydropower output and climate change.

[0010] 4. This invention applies the trained machine learning model to the hydropower generation sequence of a large-scale climate change prediction dataset obtained from the second sample of the target region; then combines bootstrap sampling to conduct significance testing, thereby quantitatively characterizing the uncertainty features of the power generation prediction results in the time and scenario dimensions, providing a robust, interpretable and engineering-application-valued method to support the quantitative assessment of uncertainty in hydropower systems by integrating climate change data. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram illustrating a hydropower prediction method under climate change that integrates machine learning methods, provided as an embodiment of this application; Figure 2 Figure showing the calibration and verification results of the SWAT model at major hydrological control stations in the upper reaches of the Yangtze River. Figure 3 This is a projection of the power generation of a large reservoir in the upper reaches of the Yangtze River under the first sample of future climate change scenarios. Figure 4 This is a schematic diagram of a decision tree model for predicting the power generation of a large reservoir in the upper reaches of the Yangtze River, based on climate characteristic values ​​and power generation under the first sample future climate scenario. Detailed Implementation

[0013] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0014] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0015] The technical solution provided in this application will now be described with reference to the accompanying drawings.

[0016] To facilitate understanding of the embodiments of this application, the following points will be explained first: First, in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the preceding and following related objects, but it does not exclude the possibility of indicating an "and" relationship; the specific meaning can be understood in context. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c; a and b; a and c; b and c; or a and b and c. Here, a, b, and c can be single or multiple.

[0017] Second, the use of prefixes such as "first" and "second" in this application is solely for the purpose of distinguishing and describing different things belonging to the same category, and does not constrain the order, size, or quantity of things. For example, "first message" and "second message" are simply different messages, and there is no chronological, size, or priority relationship between them.

[0018] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0019] like Figure 1The diagram illustrates a hydropower prediction method under climate change that integrates machine learning, as provided in this application embodiment. It represents a hydropower assessment framework that evolves from traditional physical models to data-driven and extended applications: The left side, based on simulations under the first sample condition, starts from different greenhouse gas emission scenarios and combines climate factors such as temperature and precipitation output from global climate models. After downscaling, it drives a distributed hydrological model, reservoir scheduling, and hydropower model to obtain basic results such as hydropower generation. The middle section extracts climate and hydrological output characteristics to construct hydropower output indicators and introduces a generalized machine learning model for hydropower prediction, compressing complex physical processes into a generalizable power generation mapping relationship. The right side, based on this, applies machine learning algorithms to newly added and extended climate scenarios to achieve rapid power generation prediction and further evaluate changes in hydropower output and their uncertainty range, thus forming a hydropower prediction and uncertainty analysis system that integrates climate change, hydrological processes, scheduling constraints, and intelligent models.

[0020] The method includes the following steps: A hydrological model for the target area was established. Historical runoff observation data from representative hydrological stations were used to calibrate and validate the parameters of the hydrological model, and simulation models and historical runoff simulation data that can be used to simulate and predict the impact of climate change on runoff in the target area were obtained.

[0021] The hydrological model is coupled with the reservoir scheduling model and the hydropower model. The historical runoff simulation data of representative hydrological stations in the target area are used to drive the reservoir scheduling model. The parameters of the reservoir scheduling model and the hydropower model are optimized to establish a hydrological-reservoir-hydropower coupled model.

[0022] We will obtain and analyze the first sample of climate change prediction datasets for the target region, and perform downscaling and bias correction on the future climate change prediction datasets of the first sample of the target region.

[0023] The hydrology-reservoir-hydropower coupled model is driven by the data of the first sample under different future climate scenarios to obtain the time series of power generation of the first sample in the target area under future climate scenarios.

[0024] Based on the time series of power generation under future climate scenarios of the first sample in the target area, and combined with hydropower generation and climate change characteristics, a generalized machine learning model for hydropower prediction is established based on machine learning algorithms.

[0025] We obtain a dataset of future climate change predictions for the second sample of the target region, drive a generalized machine learning model for hydropower prediction based on the climate change characteristics of the second sample, obtain the future hydropower prediction results under the second sample of the target region, and quantify the uncertainty of the future hydropower prediction sequence under the second sample of the target region.

[0026] It should be noted that the first sample refers to the basic dataset used to drive the hydrology-reservoir-hydropower coupling model and establish the generalized model of hydropower prediction machine learning, which is a finite sample. The second sample refers to the data set formed by the combination of multiple climate models, multiple emission scenarios, and multiple time series, which is a large sample.

[0027] It should be noted that climate models may exhibit systematic biases when simulating climate change in certain regions, such as discrepancies between simulated temperature or precipitation data and actual observational data. Bias correction aims to adjust the values ​​of climate variables (such as temperature, precipitation, wind speed, and humidity) output by the climate model to better align with observational data, thereby improving the reliability of predictions. Furthermore, climate models typically operate at a regional scale, resulting in low resolution and an inability to capture the details of localized areas. For regional issues such as hydropower production, agriculture, and water resource management, higher spatial resolution climate data is often required. Downscaling transforms low-resolution climate model data into high-resolution data for the target region by utilizing the statistical relationships or dynamic mechanisms between the output of large-scale climate models and local climate influencing factors, ensuring that climate predictions accurately reflect climate change in the target area.

[0028] It should be noted that the future climate change projection dataset for the first sample in the target region can be datasets released by international Coupled Model Intercomparison (CMIP) projects initiated by the World Climate Research Programme, including but not limited to CMIP5 (Fifth Coupled Model Intercomparison Project) and CMIP6 (Sixth Coupled Model Intercomparison Project) datasets. A small sample set from the acquired dataset will be downscaled to obtain training samples. Model data downscaling can employ, but is not limited to, internationally accepted statistical downscaling methods such as transformation functions, circulation patterns, and weather generators. If there is sufficient historical observation data, historical observation sequences can be used as input, eliminating the need to downscale the climate model projection data to obtain training samples. The future climate change projection dataset for the second sample in the target region can be the latest CMIP6 dataset and any new scenario datasets that may be released in the future (such as the Seventh Coupled Model Intercomparison Project or other subsequent versions).

[0029] Specific Implementation Example: The upper reaches of the Yangtze River have abundant water resources and a large elevation difference, resulting in extremely rich hydropower resources. Since the 1950s, numerous large reservoirs have been built in this region, forming the world's largest reservoir group. Under the influence of global climate change, the efficiency of these reservoir power stations faces new challenges. This implementation example takes the upper reaches of the Yangtze River as the research area and selects one reservoir as the object to establish a hydrological-reservoir-hydropower coupling model and a generalized machine learning model for hydropower prediction. The upper reaches of the Yangtze River can be divided into five major sub-basins: the Jinsha River, the Minjiang River, the Tuojiang River, the Jialing River, the Wujiang River, and the main stream section of the upper reaches of the Yangtze River. Each of the main stream and tributaries has a major control hydrological station.

[0030] The parameters of the hydrological model are calibrated and validated, and the specific process is as follows: Collect climate observation data, underlying surface characteristic parameters, and historical runoff observation data from representative hydrological stations within the target area.

[0031] Hydrological models are built and parameterized using climate observation data and underlying surface characteristic parameter data within the target area.

[0032] The hydrological model outputs historical runoff simulation results for representative hydrological stations within the target area, which are then compared with the corresponding runoff observation data. This process is used to calibrate and validate the parameters of the hydrological model, thereby obtaining historical runoff simulation data for representative hydrological stations within the target area.

[0033] In this embodiment, the historical runoff simulation results of representative hydrological stations in the target area are obtained by running the hydrological model, and compared with the corresponding runoff observation data. Accuracy evaluation index is selected, and the key parameters in the model are iteratively adjusted through parameter optimization algorithm. The model is run repeatedly until the simulation accuracy reaches the preset threshold to obtain the optimized parameter set.

[0034] It should be noted that hydrological models are important tools for simulating hydrological processes and understanding hydrological patterns. Since the 1960s, hydrologists both domestically and internationally have established and developed various hydrological models with different conceptual interpretations, such as SWAT, VIC, and the Xin'anjiang model. Extensive practical experience has demonstrated that when conducting hydrological simulation analysis in a specific watershed, an appropriate hydrological model can be selected based on the specific conditions of the watershed, combined with the modeling principles and applicability of the hydrological model.

[0035] It should be noted that the coefficient of determination (R²) is typically used as the evaluation index for the accuracy of hydrological model calibration and verification. 2 ) and Nash efficiency coefficient (E ns The calculation formula is as follows: ; ; In the formula, R 2 To determine the coefficients, Ens is the Nash efficiency coefficient, Oi As a representative hydrological station's measured runoff sequence, S i For representative hydrological stations, simulated runoff sequences, This is the average value of the measured values. To simulate the average value, i is an identifier for each set of observations in the dataset, i = 1, 2, 3, ..., N, where N is the total number of observation identifiers.

[0036] R 2 It reflects the consistency between the trends of simulated and measured values; the closer its value is to 1, the better the consistency. ns R reflects the overall efficiency of the model; the closer its value is to 1, the higher its applicability. Generally, R is considered to be... 2 >0.5, E ns When the value is greater than 0.5, the simulation effect is acceptable.

[0037] It should be noted that after parameter calibration and verification, the hydrological model can output historical runoff simulation results that match the actual situation, thus providing reliable input for subsequent reservoir scheduling simulation and power generation assessment.

[0038] like Figure 2 The figure shows the calibration and verification results of the SWAT model at major hydrological control stations in the upper reaches of the Yangtze River. Overall, after parameter calibration, the simulated runoff time series of hydrological stations at control sections of each sub-basin are basically consistent with the observed seasonal changes in runoff. Statistical analysis of accuracy evaluation indicators shows that the Ens efficiency coefficients of each sub-basin during the calibration period range from 0.85 to 0.89, and the coefficient of determination R0 is... 2 The efficiency coefficient of Ens during the validation period is between 0.87 and 0.91, and the coefficient of determination R is between 0.81 and 0.91. 2 The values ​​ranged from 0.88 to 0.92, and there was no significant difference between the calibration and validation periods, indicating that the calibrated SWAT model can be used to simulate and assess runoff impacts of climate change in the upper reaches of the Yangtze River.

[0039] The process of establishing a hydrological-reservoir-hydropower coupled model includes: Historical runoff simulation data from representative hydrological stations within the target area are imported into the reservoir scheduling model. The scheduling simulation is then performed by combining the reservoir capacity-water level relationship curves of each reservoir within the target area stored in the database, thereby obtaining the time series of outflow and reservoir water level for the target area.

[0040] The reservoir water level time series corresponding to the target area is compared with the annual key water level threshold data of reservoirs under the conventional scheduling rules stored in the database, and the scheduling rule parameters in the reservoir scheduling model are optimized.

[0041] It should be noted that the key water level threshold data of the reservoir includes, but is not limited to, water level, flood limit water level, normal storage water level, dead water level, and flood control high water level.

[0042] Meanwhile, the reservoir scheduling model will set the adjustment range of reservoir outflow within each target time step based on the predetermined optimization objectives and reservoir constraints, and select the reservoir outflow within the corresponding target time step within the set adjustment range of reservoir outflow within each target time step.

[0043] A hydropower model is established by collecting characteristic parameters such as the installed capacity of the reservoir power station, the design head of the turbine generator, and the comprehensive efficiency coefficient of the unit in the target area. The real-time outflow and water level of the reservoir in the target area are input into the hydropower model to establish a hydrological-reservoir-hydropower coupling model.

[0044] The climate change prediction dataset of the first sample in the target region is downscaled and biased. The specific process is as follows: Various downscaling processes were applied to the first sample of the climate change prediction dataset in the target region.

[0045] Extract the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method.

[0046] The root mean square error (RMSE) of the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method is calculated and averaged to obtain the mean RMSE. At the same time, the correlation coefficients between the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method and the actual observed climate dataset are calculated.

[0047] The mean root mean square error (RMSE) values ​​of the climate models corresponding to each downscaling method are sorted from smallest to largest. Based on the sorting of the RMSE values ​​output by the climate models corresponding to each downscaling method, the downscaling method with the largest correlation coefficient is selected as the first combination of climate model and downscaling method.

[0048] It should be noted that in the actual analysis process, the datasets of five different climate models under two greenhouse gas emission scenarios (RCP4.5: intermediate emission path; RCP8.5: high emission path) were first selected from the CMIP5 climate model. The specific models are HaDGEM2-ES (UK Met Office climate model), NorESM1-M (Norwegian Meteorological Service climate system model), IPSL-CM5A-LR (Pierre Simon Laplace Institute climate model), MIROC-ESM-CHEM (Japan Agency for Marine-Earth Science and Technology climate model), and GFDL-ESM2M (US Geophysical and Hydrodynamic Laboratory climate model). The data were taken from 1986 to 2005 as the baseline period and 2006 to 2100 as the projection period. The relevant climate variable output data, such as precipitation, average temperature, maximum temperature, and minimum temperature, were obtained as the first sample of the future climate change projection dataset for the target area. Subsequently, the CN05.1 gridded meteorological observation dataset, constructed by the China Meteorological Administration based on measured data from more than 2,400 meteorological observation stations across the country through scientific interpolation, was used as the validation standard to downscale these climate model data.

[0049] It should be noted that downscaling and bias correction of the first sample of climate change prediction datasets in the target region can be performed using the spatial disaggregation statistical downscaling method and the equidistant cumulative distribution function (EPF) bias correction method. Then, the annual mean temperature, annual precipitation, total number of precipitation days, and interannual variation coefficients of the climate models corresponding to each downscaling method are extracted, and the correlation coefficient and root mean square error (RMSE) with the observed climate dataset are calculated. The mean RMSE values ​​obtained from the climate models corresponding to each downscaling method are sorted from smallest to largest, and the downscaling method with the highest correlation coefficient is selected as the first combination of climate model and downscaling method based on the sorted RMSE values ​​of the climate models corresponding to each downscaling method.

[0050] The specific process for obtaining the power generation time series of the first sample in the target area under future climate scenarios is as follows: The hydrology-reservoir-hydropower coupled model is driven by climate data under various future climate scenarios, and the output is the time series of power generation of each reservoir and power station in the target area under the future climate scenarios.

[0051] It should be noted that the output of the first sample of power generation time series of each reservoir power station in the target area under future climate scenarios includes the process of the reservoir scheduling model determining the outflow within each target time step. Specifically, setting the adjustment range of reservoir outflow within each target time step follows these steps: First, obtain the real-time status of the reservoir at the current time step, mainly including the reservoir water level and inflow. Next, based on the set optimization objectives (such as maximizing power generation, ensuring water supply, or flood control safety), the target range of outflow is initially calculated. That is, within the current time step, after comprehensively considering the real-time status of the reservoir, operational rule constraints, and scheduling optimization objectives, the model considers a range of outflow values ​​that is both safe and feasible and conducive to achieving the scheduling objectives. Within this range, the model further couples multiple constraints of the reservoir (such as the allowable maximum / minimum water level, maximum outflow capacity, minimum discharge flow based on downstream demand, etc.), and solves the problem through constraint pruning and optimization algorithms to finally determine the optimal outflow for that period.

[0052] It should be noted that the hydropower model is based on the key characteristic parameters of the generator set and uses the real-time water level and outflow of the reservoir as dynamic driving inputs. Based on the rated head, power generation efficiency coefficient, and other characteristic parameters of the turbine generator set, the model dynamically calculates the net head based on the real-time water level, and then calculates the real-time power output and power generation through the head, outflow, and power generation efficiency coefficient. The calculation formula can be briefly expressed as follows: P = η × ρ × g × Q × H; Where P is the real-time output (W), η is the unit's overall efficiency coefficient, ρ is the density of water, g is the gravitational acceleration, Q is the reservoir outflow (m³ / s), and H is the real-time net head (m).

[0053] The output includes the time series of power generation of the first sample in the target area under future climate scenarios, and also includes: The power generation time series of the first sample in the target area under future climate scenarios is statistically analyzed to obtain the statistical results of the power generation time series under each climate scenario, including the long-term average power generation, maximum value, coefficient of variation, and the magnitude of power generation change in adjacent time steps. Anomaly judgment analysis is performed on the statistical results of the power generation time series under each climate scenario.

[0054] The anomaly judgment analysis of the statistical results of power generation time series under various climate scenarios includes the first anomaly and the second anomaly. The statistical results of power generation time series under various climate scenarios are corrected based on the first anomaly and the second anomaly of the statistical results of power generation time series under various climate scenarios.

[0055] It should be noted that the anomaly analysis of the statistical results of power generation time series under various climate scenarios includes first anomalies and second anomalies. The specific process is as follows: First, based on the judgment result of the first anomaly, that is, when the average power generation, maximum value, or coefficient of variation (the coefficient of variation is a statistic used to measure the volatility of power generation time series, usually the ratio of the standard deviation of power generation to its average value) in the statistical results of power generation time series under a certain climate scenario exceeds the reasonable range of the corresponding target area hydropower system under the existing installed capacity and dispatch rules in the database within a preset time window, this anomaly is considered to reflect the first anomaly. At this time, by retrospectively examining the runoff simulation results and reservoir dispatch process corresponding to this climate scenario, the key parameters affecting the upper limit of power generation and annual allocation are constrained and corrected. That is, additional constraints are added to the reservoir dispatch model, such as limiting the maximum outflow or minimum storage capacity of the reservoir, to ensure that the corrected power generation returns to a reasonable range in terms of long-term statistical characteristics. This ensures that the corrected power generation series returns to a reasonable range in terms of long-term statistical characteristics that matches the installed capacity and hydrological conditions, while maintaining the relative relationship of the original climate change trend's impact on power generation. After correcting the first anomaly, further refined adjustments are made for the second anomaly. The second anomaly primarily manifests as power generation changes in adjacent time steps exceeding or equaling the power generation change threshold stored in the database. These anomalies often do not alter the long-term average level but significantly amplify the volatility of the power generation sequence. For this type of anomaly, instead of compressing or shifting the entire sequence, local backtracking is performed on the anomalous section based on climate input changes and hydrological response characteristics corresponding to the anomaly's occurrence period. Power generation changes in the anomalous section are smoothed out. For example, if power generation fluctuations exceed the preset normal fluctuation range, the reservoir's water release strategy needs to be adjusted to control power generation. Adjusting the reservoir's water release strategy typically involves smoothing power generation fluctuations by limiting the increase or decrease in water release volume, while ensuring the reservoir's normal operation. This limits the maximum or minimum outflow to avoid large fluctuations in power generation caused by excessive or insufficient instantaneous water release, ensuring that the changes between adjacent time steps are consistent with the reservoir's scheduling and generating unit regulation capabilities, thereby eliminating non-physical sudden increases or decreases.

[0056] like Figure 3 The figure shows the projected power generation of a large reservoir in the upper reaches of the Yangtze River under the first sample future climate change scenario. Based on the hydrological-reservoir-hydropower coupling model, the power generation of a large reservoir in the upper reaches of the Yangtze River under the first sample future climate scenario is shown. It displays the annual power generation of a large reservoir in the upper reaches of the Yangtze River from 2026 to 2099 under the RCP4.5 medium emission scenario and the RCP8.5 high emission scenario, based on five different climate models. It can be seen that the power generation under different greenhouse gas emission scenarios and different climate models predicting future climate change fluctuates significantly within a certain range.

[0057] The generalized machine learning model for hydropower prediction is established based on machine learning algorithms. The specific process is as follows: The target region is used as the training dataset by combining the climate model and the downscaling method under the first sample future climate scenario, along with the power generation under each corresponding climate scenario. Feature parameters are then extracted from the training dataset.

[0058] A generalized machine learning model for predicting hydropower generation in a target area is established using a decision tree model, which relates the hydropower generation of reservoirs to the characteristic parameters of climate factors.

[0059] It should be noted that the feature parameters of the training dataset include the annual average, interannual coefficient of variation, and seasonal distribution concentration of key climate factors such as precipitation and average temperature. The training dataset is calculated and generated from the climate data corresponding to the first combination of climate models and downscaling methods in the target region, as well as the power generation simulation results under various climate scenarios in the target region. The specific process can be understood as statistical analysis and feature engineering of the climate variable series to ensure that the decision tree model can capture the impact of climate change on hydropower. First, the climate data in the training dataset needs to be organized by time series, such as daily, monthly, or seasonal temperature and precipitation data, and aligned with the corresponding power generation data for each time step to ensure a one-to-one correspondence between the climate input and the reservoir outflow or power generation output at each time step. Based on this, annual-scale statistical analysis can be performed on each climate variable. The annual average is calculated by taking the arithmetic mean of all time steps of the climate variable within a year, reflecting its overall level. The interannual coefficient of variation is calculated by the ratio of the standard deviation to the mean of the annual average series of consecutive years, reflecting the fluctuation range or instability of the climate variable between different years. The seasonal distribution concentration is calculated by dividing the year into several seasons (e.g., spring, summer, autumn, winter) or by month, calculating the proportion of the climate variable in each season or month in the total annual amount, and further quantifying the degree of concentration of its seasonal distribution through concentration indices (such as standardized entropy values ​​or concentration coefficients). After feature extraction, each time series is mapped to one or more feature vectors. These vectors are matched with power generation data in the same period or year to form the final training samples. Each sample record contains a climate feature vector and the corresponding power generation output. The decision tree model can learn the influence pattern of climate factors on hydropower during the training phase, including long-term trends, interannual fluctuations and seasonal distribution characteristics. Thus, when applied to new future climate scenarios, it can predict power generation based on the key climate factor characteristics input. It also provides basic data and interpretable feature inputs for subsequent multi-scenario and multi-mode uncertainty analysis.

[0060] like Figure 4The diagram shows a decision tree model for predicting the power generation of a large reservoir in the upper reaches of the Yangtze River, based on climate characteristic values ​​and power generation under the first sample future climate scenario. The decision tree model uses power generation as the response variable and constructs decision rules using annual cumulative precipitation (r_sum), precipitation concentration (r_pcd), precipitation variation coefficient (r_sd), maximum monthly precipitation (r_max), average temperature (t_mean), and maximum temperature (t_max) as characteristic variables. The power generation value corresponding to the root node is 101 TWh, covering all 2163 samples (accounting for 100%). The model first branches based on whether the statistical measure r_sum is less than 831 mm. If the condition is met, it enters the left subtree and then branches based on whether r_sum is less than thresholds such as 800, 769, 748, 705, and 643 mm. It further subdivides based on r_max (e.g., r_max < 141) until the bottom leaf node. Each node corresponds to a predicted power generation value (e.g., power = 71) and is labeled with the number of samples and its percentage in the whole set (e.g., n = 58, 3%). If the root node does not meet the condition r_sum < 831 mm, it enters the right subtree and further combines statistical indicators such as t_mean (e.g., ≥ 8.3) and r_pcd (e.g., < 0.43) to form subdivision rules. By combining the thresholds of various climate variables, the model finally achieves the prediction of power generation.

[0061] The specific process for obtaining the future hydropower prediction results for the second sample of the target area is as follows: The climate change forecast data of the second sample of the target area under future climate scenarios are imported into the generalized machine learning model for hydropower forecasting between power generation and climate factor characteristic parameters. The power generation of the second sample of the target area is then predicted, and the power generation forecast results under each climate scenario of the target area are obtained.

[0062] The uncertainty of the future hydropower prediction sequence under the second sample of the target area is quantified. The specific process is as follows: The project extracts the future hydropower projection results under the second sample of the target area, and statistically analyzes the ensemble mean and standard deviation of the power generation projection results of each climate model for typical future periods under the second sample of the target area. At the same time, it obtains the significance p-value of the power generation change between the projection period and the measured period based on the Bootstrap sampling method. By combining the standard deviation of power generation within the projection period of the target area and the p-value of the power generation change trend between the projection period and the measured period, the project jointly conducts significance tests and quantifies the uncertainty of the power generation projection results under each climate scenario in the target area.

[0063] It should be noted that the significance test of the change trend of power generation in the future estimated period relative to the baseline period based on Bootstrap sampling refers to the process of repeatedly drawing samples from the existing data with replacement to generate multiple new sample sets. Under the premise of not relying on the assumption that the data strictly follows a normal distribution, an empirical distribution of the statistic is constructed, and the reliability of the change trend in the future period relative to the baseline period is estimated accordingly. Specifically, firstly, for any climate model, its power generation time series for the baseline period and the future projected period are extracted. Bootstrap resampling with replacement is performed on the annual power generation series for both periods (e.g., 1000 times). Each resampling generates a new pair of baseline and future period samples. For each pair of bootstrap samples, the mean power generation is calculated, and the difference between the future period mean and the baseline period mean is recorded. By constructing the bootstrap empirical distribution of this difference (or the corresponding t-statistic), the significance p-value of the predicted power generation change trend of the climate model is calculated. If the p-value is less than the preset significance level (e.g., α=0.05), the power generation change trend shown by the model is considered statistically significant. Secondly, after all participating climate models have completed the above independent tests, the number of models that pass the significance test (i.e., p<0.05) is counted. If the number of models that pass the significance test exceeds half of the total number of models (i.e., >50%), the power generation change trend in the future period relative to the baseline period is considered statistically significant.

[0064] Table 1 shows the statistical values ​​of power generation changes of a large reservoir in the upper reaches of the Yangtze River based on the second sample future climate scenario. The second sample climate change prediction dataset was constructed using predictions from 14 CMIP6 climate models under three scenarios: SSP1-2.6 (low emissions), SSP2-4.5 (medium emissions), and SSP5-8.5 (high emissions). The baseline period is 1961-2000, and the prediction periods are 2021-2060 and 2061-2100, respectively. According to the data in Table 1: under the SSP1-2.6 scenario, the power generation in 2021-2060 and 2061-2100 compared to the baseline period... The change rates for the two time periods were 0.8% (uncertainty interval ±5.6%) and 1.2% (uncertainty interval ±5.7%), respectively, with the proportions of patterns passing the significance test being 42.9% and 35.7%, respectively, and the change trends were not significant. Under the SSP2-4.5 scenario, the change rates for the two time periods were 2.4% (±3.1%) and 5.6% (±5.4%), respectively, with the proportions of significant patterns being 71.4% and 92.9%, and the change trends were significant. Under the SSP5-8.5 scenario, the change rates were 2.9% (±3.4%) and 6.2% (±4.5%), respectively, with the proportions of significant patterns being 85.7% and 100%, and the change trends were also significant.

[0065] Table 1. Projected hydropower changes and uncertainty ranges for a large reservoir in the upper reaches of the Yangtze River under future climate scenarios in the second sample. The various features and processes described above can be used independently of each other or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Furthermore, certain method or process blocks may be omitted in some embodiments. The methods and processes described herein are not limited to any particular order, and the blocks or states associated with them may be performed in other suitable orders. For example, the described blocks or states may be performed in an order different from the order specifically disclosed, or multiple blocks or states may be combined in a single block or state. Example blocks or states may be performed serially, in parallel, or in some other manner. Blocks or states may be added to or removed from the disclosed example embodiments. The exemplary cases and components described herein may be configured differently from those described. For example, elements may be added to, removed from, or rearranged compared to the disclosed example embodiments.

[0066] The various operations of the example methods described herein can be performed at least in part by an algorithm. This algorithm can be contained in program code or instructions stored in memory (e.g., the aforementioned non-transitory computer-readable storage medium). Such an algorithm may include a machine learning algorithm. In some embodiments, the machine learning algorithm may not be explicitly programmed into the computer to perform the function, but can learn from training data to create a predictive model that performs the function.

[0067] The various operations of the example methods described herein can be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute the engine of a processor implementation that operates to perform one or more of the operations or functions described herein.

[0068] Similarly, the methods described herein can be implemented at least in part by a processor, where one or more specific processors are examples of hardware. For example, at least some operations of a method can be performed by one or more processors or an engine implemented by a processor. Furthermore, one or more processors can also be operated to support the performance of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations can be performed by a set of computers (as an example of a machine including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application programming interfaces (APIs)).

[0069] The performance of certain operations can be distributed across processors, residing not only within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented engine may reside in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example embodiments, the processor or processor-implemented engine may be distributed across multiple geographic locations.

[0070] In this specification, multiple instances may implement components, operations, or structures described as single instances. Although individual operations of one or more methods are shown and described as separate operations, one or more of the separate operations may be performed simultaneously and do not need to be performed in the order shown. Structures and functions presented as separate components in the example configuration may be implemented as composite structures or components. Similarly, structures and functions presented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document.

[0071] While an overview of the subject matter has been described with reference to specific example embodiments, various modifications and changes can be made to these embodiments without departing from the broader scope of embodiments of this disclosure. Such embodiments of the subject matter are referred to herein, individually or collectively, by the term "invention," and are used for convenience only and are not intended to limit the scope of this application to any single disclosure or concept, should more than one disclosure or concept be disclosed in fact.

[0072] The embodiments described herein have been described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Therefore, the detailed description should not be construed as limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.

Claims

1. A method for predicting hydropower under climate change by integrating machine learning, characterized in that, Includes the following steps: A hydrological model for the target area was established. Historical runoff observation data from representative hydrological stations were used to calibrate and validate the parameters of the hydrological model, and simulation models and historical runoff simulation data that can be used to simulate and predict the impact of climate change on runoff in the target area were obtained. The hydrological model is coupled with the reservoir scheduling model and the hydropower model. The historical runoff simulation data of representative hydrological stations in the target area are used to drive the reservoir scheduling model. The parameters of the reservoir scheduling model and the hydropower model are optimized to establish a hydrological-reservoir-hydropower coupled model. We will obtain and analyze the first sample of climate change prediction datasets for the target region, and perform downscaling and bias correction on the first sample of climate change prediction datasets for the target region. The hydrology-reservoir-hydropower coupled model is driven by the data of the first sample under different future climate scenarios to obtain the time series of power generation of the first sample in the target area under future climate scenarios. Based on the time series of power generation of the first sample in the target area under future climate scenarios, combined with hydropower generation and climate change characteristics, a generalized machine learning model for hydropower prediction is established based on machine learning algorithms. We obtain a dataset of future climate change predictions for the second sample of the target region, drive a generalized machine learning model for hydropower prediction based on the climate change characteristics of the second sample, obtain the future hydropower prediction results under the second sample of the target region, and quantify the uncertainty of the future hydropower prediction sequence under the second sample of the target region.

2. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The specific process for calibrating and validating the parameters of the hydrological model is as follows: Collect climate observation data, underlying surface characteristic parameters, and historical runoff observation data from representative hydrological stations within the target area; Hydrological models are built and parameterized using climate observation data and underlying surface characteristic parameter data within the target area; The hydrological model outputs historical runoff simulation results for representative hydrological stations within the target area, which are then compared with the corresponding runoff observation data. This process is used to calibrate and validate the parameters of the hydrological model, thereby obtaining historical runoff simulation data for representative hydrological stations within the target area.

3. The hydropower prediction method based on climate change incorporating machine learning as described in claim 2, characterized in that: The establishment of the hydrological-reservoir-hydropower coupling model includes: Historical runoff simulation data from representative hydrological stations within the target area are imported into the reservoir scheduling model. The scheduling simulation is then performed by combining the reservoir capacity-water level relationship curves of each reservoir within the target area stored in the database, thereby obtaining the outflow and reservoir water level time series for each reservoir in the target area. The reservoir water level time series corresponding to the target area is compared with the annual key water level threshold data of reservoirs under the conventional scheduling rules stored in the database, and the scheduling rule parameters in the reservoir scheduling model are optimized. Meanwhile, the reservoir scheduling model will set the adjustment range of reservoir outflow within each target time step based on the predetermined optimization objectives and reservoir constraints, and select the reservoir outflow within the corresponding target time step within the adjustment range of reservoir outflow within each target time step. A hydropower model is established by collecting characteristic parameters such as the installed capacity of the reservoir power station, the design head of the turbine generator, and the comprehensive efficiency coefficient of the unit in the target area. The real-time outflow and water level of the reservoir in the target area are input into the hydropower model to establish a hydrological-reservoir-hydropower coupling model.

4. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The process of analyzing the climate change prediction dataset of the first sample in the target area is as follows: The climate change prediction dataset of the first sample in the target area is analyzed. Relevant climate variable output data are obtained with a preset baseline period, including annual average temperature, annual precipitation, total number of precipitation days and interannual variation coefficient. At the same time, the actual observed climate dataset is used as a validation standard to quantitatively evaluate the prediction dataset and analyze the impact of various climate models on the simulation effect of climate characteristics.

5. The hydropower prediction method based on climate change incorporating machine learning as described in claim 4, characterized in that: The specific process of downscaling and bias correction of the climate change prediction dataset of the first sample in the target region is as follows: Various downscaling processes were applied to the first sample of the climate change prediction dataset in the target region. Extract the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method; The root mean square error (RMSE) of the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method is calculated and averaged to obtain the mean RMSE. At the same time, the correlation coefficients of the annual average temperature, annual precipitation, total number of precipitation days, and interannual variation coefficient of the climate models corresponding to each downscaling method with the actual observed climate dataset are calculated. The mean root mean square error (RMSE) values ​​of the climate models corresponding to each downscaling method are sorted from smallest to largest. Based on the sorting of the RMSE values ​​output by the climate models corresponding to each downscaling method, the downscaling method with the largest correlation coefficient is selected as the first combination of climate model and downscaling method.

6. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The specific process for obtaining the power generation time series of the first sample in the target area under future climate scenarios is as follows: The hydrology-reservoir-hydropower coupled model is driven by climate data under various future climate scenarios, and the output is the time series of power generation of each reservoir and power station in the target area under the future climate scenarios.

7. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The process of obtaining the power generation time series of the first sample in the target area under future climate scenarios also includes: The power generation time series of the first sample in the target area under future climate scenarios is statistically analyzed to obtain the statistical results of the power generation time series under each climate scenario, including the long-term average power generation, maximum value, coefficient of variation, and the magnitude of power generation change in adjacent time steps. Anomaly judgment analysis is performed on the statistical results of the power generation time series under each climate scenario. The anomaly judgment analysis of the statistical results of power generation time series under various climate scenarios includes the first anomaly and the second anomaly. The statistical results of power generation time series under various climate scenarios are corrected based on the first anomaly and the second anomaly of the statistical results of power generation time series under various climate scenarios.

8. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The specific process of establishing a generalized machine learning model for hydropower prediction based on machine learning algorithms is as follows: The first combination of the climate model and downscaling method in the target area, as well as the power generation under each climate scenario in the target area, are used as the training dataset, and the feature parameters in the training dataset are extracted. The characteristic parameters include the annual average value, interannual coefficient of variation, and seasonal distribution concentration of key climate factors; A generalized machine learning model for predicting hydropower generation between reservoir power stations and climate factor characteristic parameters within a target area is established using a decision tree model.

9. The hydropower prediction method based on climate change incorporating machine learning as described in claim 1, characterized in that: The specific process for obtaining the future hydropower prediction results under the second sample of the target area is as follows: The future climate change forecast data of the second sample in the target area are imported into the generalized machine learning model for hydropower forecasting between power generation and climate factor characteristic parameters in the target area. The power generation of the second sample in the target area is then forecasted to obtain the future hydropower forecast results under the second sample in the target area.

10. The method for predicting hydropower under climate change by incorporating machine learning as described in claim 1, characterized in that: The uncertainty of the future hydropower prediction sequence under the second sample of the target area is quantified, and the specific process is as follows: The multi-model future hydropower projection results of the second sample in the target area are extracted. The standard deviation of the multi-model power generation within the projection period in the second sample of the target area is used as the uncertainty interval. At the same time, the p-value of the power generation change trend between the projection period and the measured period is obtained based on the bootstrap sampling method. The significance of the power generation projection results under various climate scenarios in the target area is tested by combining the standard deviation of power generation within the projection period in the target area and the p-value of the power generation change trend between the projection period and the baseline period.

Citation Information

Patent Citations

  • CN120562799A