Method for extracting reservoir group operation rules based on runoff space-time two-dimensional random simulation
By using two-dimensional spatiotemporal stochastic simulation of runoff and data mining machine learning methods, multi-reservoir runoff data is generated, reservoir group scheduling is optimized, the problem of insufficient spatiotemporal correlation of runoff is solved, and the scientificity and applicability of reservoir group scheduling rules are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA INST OF WATER RESOURCES & HYDROPOWER RES
- Filing Date
- 2025-12-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are unable to effectively address the spatiotemporal two-dimensional correlation of runoff, cannot generate runoff data covering massive scenarios, lack intelligent scheduling rule models, and are insufficient in the design and response to extreme scenarios.
A two-dimensional stochastic simulation method for runoff in spatiotemporal space was adopted. By screening the runoff marginal distribution and the Copula joint distribution function, and combining the Monte Carlo method, multiple sets of runoff sequence data were generated. The scheduling model was optimized using a genetic algorithm, and the scheduling rules of the reservoir group were established by combining data mining and machine learning.
It has improved the scientific and intelligent level of reservoir group scheduling rules, making them applicable to both normal and extreme conditions, and providing a reliable data foundation and scheduling decision support.
Smart Images

Figure CN121581569B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of reservoir group optimization scheduling, specifically involving a method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal random simulation of runoff. Background Technology
[0002] Reservoir projects offer comprehensive benefits, including flood control, water supply, power generation, ecological conservation, and navigation. Reservoir groups within the same system are interconnected in terms of hydraulics, electricity, and hydrology; joint scheduling can maximize social, economic, and ecological benefits. To optimize the joint operation of reservoir groups, numerous optimization methods have been developed theoretically over the past few decades, including linear programming (LP), nonlinear programming (NLP), dynamic programming (DP), and intelligent optimization methods. However, in practice, the theoretical achievements of reservoir optimization scheduling are difficult to apply directly, still relying on the experience of dispatchers combined with reservoir scheduling rules.
[0003] Reservoir scheduling rules are crucial guidelines for reservoir operation in practice. Traditional reservoir scheduling rules primarily take the form of scheduling diagrams and scheduling functions. These methods are relatively intuitive and easy to operate, hence their widespread application. Reservoir scheduling diagrams typically use time as the horizontal axis and water level as the vertical axis, dividing the map into different operational zones through key control lines. These zones correspond to differentiated scheduling strategies, helping managers make scientific decisions based on real-time water conditions to maximize the reservoir's overall benefits and minimize risks. Scheduling functions, on the other hand, describe scheduling strategies using pre-defined explicit functions (such as multiple linear regression). By comparing different function forms multiple times, a mapping relationship is established between reservoir scheduling decision variables (average outflow in the current period / reservoir capacity at the end of the current period) and decision factors (inflow in the current period and several periods before and after, initial reservoir capacity in the current period and several periods before, outflow in the previous period, etc.) to guide reservoir operation.
[0004] In recent years, with the rapid development of artificial intelligence technology, data-driven research methods based on data mining and machine learning (Artificial Neural Networks (ANN), Support Vector Machines (SVM), Random Forests (RF), Long Short-Term Memory (LSTM), etc.) to extract reservoir scheduling rules to guide reservoir operation have emerged due to their ability to effectively address the randomness, nonlinearity, and multivariate nature of reservoir scheduling decisions. Data mining techniques can be used to analyze and establish the correlation between decision variables and various decision factors, thus laying the foundation for machine learning methods to extract scheduling rules. This method does not require a pre-defined fixed function form; instead, it leverages the algorithm's powerful mapping capabilities to learn from input data, identify potential scheduling patterns, and automatically generate adaptive scheduling strategies by adjusting the model's network structure and parameters. Compared to traditional scheduling methods, reservoir scheduling rule extraction based on data mining and machine learning exhibits greater flexibility and adaptability.
[0005] The limitations of existing technologies are mainly reflected in the following aspects: First, traditional reservoir scheduling rules are generally established using historical design runoff data, making it difficult to respond promptly to runoff changes caused by climate change and human activities, i.e., the uncertainty and randomness of runoff. For example, the current design scheduling chart of Longyangxia Reservoir was compiled in 1998, and changes in the runoff sequence have caused the scheduling rules based on the assumption of hydrological consistency to become disconnected from the actual reservoir scheduling. Second, traditional reservoir scheduling rules are usually established using explicit function forms such as multiple linear regression, which are difficult to effectively reflect the nonlinear and multivariate characteristics of reservoir scheduling decisions. Third, the establishment of reservoir scheduling rules using data mining combined with machine learning methods usually relies on long-term hydrological data. However, due to the short monitoring period of some hydrological stations, it is difficult to encompass diverse runoff scenarios. Although some studies have carried out hydrological stochastic simulation research to address the problem of insufficient hydrological sequence length, these studies mostly focus on single reservoirs or single stations, emphasizing temporal correlation and insufficiently considering the spatial connections of inflows into reservoir groups. Currently, few studies have established runoff sequence reconstruction models based on the temporal and spatial two-dimensional correlation characteristics of runoff to simulate multi-temporal variations in inflow runoff into reservoir groups. Furthermore, data mining and machine learning methods are used to establish reservoir group scheduling rules with hydraulic and electrical connections under multi-temporal runoff encounter scenarios. Fourthly, with climate change and increased human activities, extreme scenarios such as extremely abundant and extremely scarce runoff are becoming more frequent and intense. Existing stochastic runoff simulation models and reservoir scheduling rule extraction models mostly focus on conventional scenarios, and there are still some shortcomings in the design and response to extreme scenarios. Summary of the Invention
[0006] To address the aforementioned shortcomings in existing technologies, the present invention provides a method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff. This method solves the problems in existing methods where stochastic simulation of runoff fails to fully characterize the two-dimensional spatiotemporal correlation, thus failing to generate runoff data covering a vast amount of scenarios, and lacks the ability to automatically identify key decision factors and construct intelligent scheduling rule models using massive optimization scheduling results.
[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0008] A method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff is provided, which includes the following steps:
[0009] S1. Taking a group of reservoirs with hydraulic and electrical connections as the research object, we screen the runoff margin distribution function for each reservoir and each time period, and screen the runoff Copula joint distribution function for the same reservoir and different time periods, as well as for different reservoirs in the same time period, based on the runoff margin distribution function.
[0010] S2. Considering the two-dimensional spatiotemporal characteristics of runoff, using the selected runoff Copula joint distribution function and combining the inverse transformation sampling of the Monte Carlo method, multiple sets of runoff sequence data containing all reservoirs and all time periods are generated through multiple random simulations.
[0011] S3. Determine the objective function and constraints for reservoir group optimization, establish a reservoir group optimization scheduling model, use multiple sets of runoff sequence data from multiple random simulations as model input, use optimization algorithms to solve the model, and obtain the corresponding multiple sets of reservoir group optimization scheduling results, including all reservoirs, the end-of-period reservoir capacity for all time periods, the average outflow for each time period, the average power generation flow for each time period, the average water discharge for each time period, and the average power output for each time period.
[0012] S4. Based on multiple sets of runoff sequence data from multiple random simulations and their corresponding multiple sets of reservoir group optimization scheduling calculation results, data mining methods are used to determine key decision factors, and machine learning methods are used to establish a reservoir group scheduling rule model.
[0013] S5. Using the state variables of the current and several preceding / following time periods of the determined key decision factors as input to the reservoir group scheduling rule model, the scheduling operation results of the reservoir group are predicted. The state variables include the position of the time period in a year and the measured and predicted inflow of all reservoirs in the current and several preceding / following time periods, the measured initial reservoir capacity in the current and several preceding time periods, and the measured average outflow in the several preceding time periods. The scheduling operation results include the final reservoir capacity of all reservoirs in the current and several preceding time periods, the average outflow, the average power generation flow, the average water discharge flow, and the average power output.
[0014] The beneficial effects of the above technical solution are as follows: This solution proposes a method system for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff, realizing a complete technical path of "two-dimensional spatiotemporal correlation analysis of runoff - stochastic simulation of runoff in multiple reservoirs - optimized scheduling of reservoir groups - data mining of key decision factors - extraction of scheduling rules through machine learning".
[0015] Furthermore, step S1 further includes:
[0016] S11. Select multiple probability distribution functions, including at least the Normal function, Gamma function, Log-normal function, GEV function and Weibull function, and calculate the runoff marginal distribution function for each reservoir and each time period respectively;
[0017] S12. The parameters of the runoff marginal distribution function are calculated using the maximum likelihood method, and the root mean square error is used as the evaluation index to test the results of various distribution functions and select the most suitable runoff marginal distribution function.
[0018] S13. Based on the runoff edge distribution function, select multiple runoff Copula joint distribution functions, including at least the Clayton Copula function, the Frank Copula function, and the Gumbel Copula function, and calculate the runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs respectively.
[0019] S14. The parameters of the runoff Copula joint distribution function are calculated using the maximum likelihood method. The root mean square error is used as the evaluation index to verify the results of various runoff Copula joint distribution functions and select the most suitable runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs.
[0020] Furthermore, given the runoff of the reservoir during time period t in time i, we can find the expression for the runoff of the reservoir during time period t+1 in time i. for:
[0021]
[0022] in, Let be the joint Copula distribution function of runoff for any reservoir i at time t and time t+1; and Let be the runoff of any reservoir i at time t and time t+1, respectively;
[0023] Given the runoff of reservoir i in time period t, find the expression for the runoff of reservoir i+1 in time period t. for:
[0024]
[0025] in, Let be the joint Copula distribution function of runoff from reservoir i and reservoir i+1 at any time interval t; and Let t represent the runoff of reservoir i and reservoir i+1 at any given time period t.
[0026] The beneficial effects of the above technical solution are as follows: This solution can select the most suitable runoff marginal distribution function for each reservoir and each time period by comparing and selecting multiple marginal distribution function forms; it can select the most suitable runoff Copula joint distribution function for the same reservoir and different time periods, as well as for the same time period and different reservoirs; it avoids the error caused by the inaccuracy of using only one marginal distribution function and one joint distribution function to describe the statistical characteristics of runoff, and provides support for ensuring the reliability of subsequent runoff stochastic simulation; it establishes runoff joint distribution functions for the same reservoir and different time periods, as well as for the same time period and different reservoirs, respectively, taking into account the correlation between adjacent time periods of any reservoir and between adjacent reservoirs in any time period, that is, it considers the two-dimensional temporal and spatial correlation characteristics of runoff.
[0027] Furthermore, step S2 includes generating multiple sets of conventional scenario runoff sequence data containing all reservoirs and all time periods. The detailed implementation method includes:
[0028] S21. To randomly simulate the runoff of Reservoir 1 in adjacent time periods, a linear congruent random number generator is used to generate random numbers a1∈(0,1); the inverse function of the cumulative distribution function in the marginal distribution is used to map the random numbers to random variables of the target distribution, i.e., to solve... The runoff x of reservoir 1 during time period 1 can be obtained. 1,1 Generate random numbers a2, a3, ..., a T Given (0,1), we can obtain the following system of equations:
[0029]
[0030] Solving the above system of equations will yield the runoff (x) of the reservoir during time period 2 to T. 1,2 , x 1,3 , …, x 1,T-1 , x 1,T );
[0031] S22. To randomly simulate the runoff of Reservoir 2 in adjacent time periods, a linear congruent random number generator is used to generate random numbers b1∈(0,1); let Given the runoff x of reservoir 1 during time period 1. 1,1 Solving this equation will yield the runoff x during time period 1 of reservoir 2. 2,1 Generate random numbers b2, b3, ..., b T Given (0,1), we can obtain the following system of equations:
[0032]
[0033] Solving the above system of equations will yield the runoff (x) of the reservoir during time period 2~T. 2,2, x 2,3 , …, x 2,T-1 , x 2,T );
[0034] S23. Following S22, perform runoff simulations for reservoirs 3 through I in sequence.
[0035] The beneficial effects of the above technical solution are as follows: This solution considers the two-dimensional correlation characteristics of runoff in time and space, and adopts the Monte Carlo inverse transform sampling method. It can generate massive amounts of multi-reservoir runoff data with temporal and spatial correlation through random simulation, effectively making up for the problem of limited time length of historical runoff monitoring data, improving the representativeness of runoff data in multi-reservoir optimization problems, and is applicable to different spatial topological connections of reservoir groups such as series, parallel, and mixed connections. It provides a more scientific and reliable data foundation for the establishment of scheduling rules for reservoir groups with hydraulic and power connections.
[0036] Furthermore, step S2 also includes generating multiple sets of extreme scenario runoff sequence data containing all reservoirs and all time periods, the implementation method of which further includes:
[0037] Frequency statistical analysis was performed on historical runoff data of all reservoirs and all time periods to define the extreme runoff judgment threshold for any reservoir and any time period. Then, using this threshold as the discrimination criterion, multiple sets of long-sequence simulated runoff of multiple reservoirs under normal scenarios were identified and labeled, and extreme scenario multi-reservoir runoff sequence data considering spatiotemporal two-dimensional correlation were obtained.
[0038] Based on the runoff sequence data of multiple reservoirs under extreme scenarios, time-level cluster analysis was used to obtain single-period extreme scenarios and multi-period extreme scenarios, and spatial-level cluster analysis was used to obtain single-reservoir extreme scenarios and multi-reservoir extreme scenarios.
[0039] The single-period extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff in any time period reaches the extreme threshold, it is defined as a single-period extreme abundance or a single-period extreme drought.
[0040] The multi-period extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff in more than one time period reaches the extreme threshold, it is defined as multi-period extreme abundance, multi-period extreme scarcity, or multi-period abundance-scarcity transition.
[0041] The extreme scenario for a single reservoir is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff of any reservoir reaches the extreme threshold, it is defined as an extreme abundance of a single reservoir, an extreme drought of a single reservoir, or a shift between abundance and drought in a single reservoir.
[0042] The multi-reservoir extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff of more than one reservoir reaches the extreme threshold, it is defined as multi-reservoir extreme abundance, multi-reservoir extreme drought, or multi-reservoir abundance-dry transition.
[0043] The beneficial effects of the above technical solution are as follows: This solution not only proposes a method for constructing runoff under conventional scenarios, but also a method for constructing runoff under extreme scenarios. Based on the long-sequence simulated runoff of multiple reservoirs under conventional scenarios, through the identification of extreme runoff and spatiotemporal clustering analysis, it further filters out different types of long-sequence simulated runoff of multiple reservoirs under extreme scenarios that consider the two-dimensional correlation between spatiotemporal factors, thus ensuring the applicability of reservoir group scheduling rules under both conventional and extreme conditions.
[0044] Furthermore, the objective function for optimizing the reservoir group is to maximize power generation, and its expression is:
[0045] ;
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] in, The objective is to maximize the power generation of a group of reservoirs with both hydroelectric and electrical connections; i is the reservoir index. I represents the total number of reservoirs in the reservoir group; t represents the time period index. T represents the total number of time periods; For the power station of reservoir i during the time period The power generation capacity; The time period is long; The density of water; It is the acceleration due to gravity; Let be the power generation efficiency of the power station in reservoir i; , and These represent the outflow from reservoir i, the power generation flow from the power station, and the wastewater discharge during time period t, respectively. , , and These represent the average head, average water level in front of the dam, average water level behind the dam, and average head loss of the power station at reservoir i during time period t. Let i be the average reservoir capacity of reservoir i during time period t; and These represent the reservoir capacity of reservoir i at the beginning and end of time period t, respectively. , , , and and , , , and All coefficients are obtained using polynomial fitting;
[0052] The constraints include water balance equations, initial / terminal reservoir capacity constraints, reservoir capacity constraints, outflow constraints, power generation flow constraints, and hydropower station output constraints.
[0053] The expression for the water balance equation is:
[0054] ;
[0055] The expression for the initial / termination storage capacity constraint is:
[0056] ; ;
[0057] The expression for the reservoir capacity constraint is:
[0058] ;
[0059] The expression for the outflow constraint is:
[0060] ;
[0061] The expression for the power generation flow constraint is:
[0062] ;
[0063] The expression for the power output constraint of the hydropower station is:
[0064] ;
[0065] in, For reservoir During the period Inbound traffic; For reservoir During the period Evaporation rate; For reservoir Storage capacity at the beginning of the scheduling period; For reservoir The target storage capacity at the end of the scheduling period; and Reservoirs During the period Minimum and maximum storage capacity; and Reservoirs During the period The minimum and maximum discharge flow rates; and Reservoirs The power station during the period The minimum and maximum power generation flow rates; and Reservoirs The power station during the period The minimum and maximum power generation capacity.
[0066] Furthermore, the optimization algorithm is a genetic algorithm, and its detailed implementation method includes:
[0067] A1. Initialize the parameters of the genetic algorithm, using multiple sets of runoff sequence data from all reservoirs and all time periods from multiple random simulations in step S2 as model inputs (normal scenario or extreme scenario).
[0068] A2. Using the total discharge flow of all reservoirs, the average discharge flow over all time periods, or the reservoir capacity at the end of a time period as decision variables, randomly generate N different decision variables as the initial population. ;
[0069] A3. Calculate the fitness of each individual in the population based on the objective function of the reservoir group optimization scheduling model;
[0070] A4. Use the tournament selection algorithm to select individuals from the population to generate a population of size N / 2. ;
[0071] A5. According to the preset crossover probability, select from the population Individuals are selected for crossover to generate offspring, and these offspring replace the original population. The parent individual in the middle;
[0072] A6. Adjust the population according to the preset mutation probability. Individuals undergo mutation operations, and the mutated individuals replace the original individuals in the next generation of the population. Chromosomes that have not undergone mutation directly enter the next generation of the population. ;
[0073] A7. Calculate the population based on the objective function of the reservoir group optimal scheduling model. The fitness value of each individual is calculated, and the current generation number Gen is incremented by 1;
[0074] A8. Determine whether the number of generations Gen exceeds the set maximum number of generations or whether all fitness values meet the set threshold. If yes, proceed to step A9; otherwise, return to step A3.
[0075] A9. Select the individual with the best fitness in the population as the optimal decision variable, and output the operating parameters corresponding to the optimal decision variable, such as the reservoir capacity at the end of the time period, the average discharge flow, the average power generation flow, the average water abandonment flow, and the average power output, for all reservoirs and all time periods. This is the optimized scheduling result.
[0076] The beneficial effects of the above technical solution are as follows: by optimizing the calculation with globality and robustness of genetic algorithm, the spatiotemporal redistribution of water volume, head and flow by reservoir regulation is fully utilized. Under the premise of meeting the constraints of water balance, water level, flow and output, the overall power generation benefits of the reservoir group and the efficient utilization of water resources are maximized.
[0077] Furthermore, step S4 further includes:
[0078] A1. Constructing a multi-scheme sample set for data mining:
[0079] Using multiple sets of runoff series data from multiple random simulations and corresponding multiple sets of reservoir group optimization scheduling calculation results, a multi-scheme sample set for data mining is constructed. Each scheme sample includes all reservoirs, initial reservoir capacity for all time periods, final reservoir capacity for all time periods, average inflow for each time period, average outflow for each time period, average power generation for each time period, average water discharge for each time period, and average power output for each time period.
[0080] A2. Use data mining methods to identify key decision factors:
[0081] B1. Calculate the grey relational degree of each decision factor.
[0082] The average outflow or end-of-period reservoir capacity of all reservoirs is selected as the decision variable. The position of the period in the year, as well as the inflow, initial reservoir capacity, and average outflow of several previous and current periods of all reservoirs, are selected as candidate decision factors. The grey relational degree between the candidate decision factors and the decision variables is calculated one by one. The specific expression is as follows:
[0083]
[0084] in, The grey relational degree of candidate decision factor j; The grey relational coefficient of candidate decision factor j in record n; [0,N], where N is the total number of records; The value of the decision variable at record n; Let j be the value of candidate decision factor j at record n; Here, the resolution coefficient is denoted by ; where, the grey relational degree is denoted by . The value range is [0,1]. The closer the value is to 1, the stronger the correlation between the two variables; the closer the value is to 0, the weaker the correlation between the two variables.
[0085] B2. Set the gray relational degree threshold Choose to satisfy The factors are used as key decision factors;
[0086] A3. Construct a multi-scheme sample set for the machine learning model based on decision variables and key decision factors, and divide the multi-scheme sample set into a training set and a test set according to a preset ratio. Use the training set and the test set to train and validate the machine learning model to obtain the reservoir group scheduling rule model.
[0087] The verification methods include using evaluation indicators such as root mean square error, mean absolute error, and coefficient of determination to evaluate the effectiveness of the reservoir group scheduling rules of the machine learning method, and using evaluation indicators such as reservoir capacity constraint over-limit rate and outflow over-limit rate to evaluate the feasibility of the reservoir group scheduling rules of the machine learning method.
[0088] Furthermore, the machine learning model is the LightGBM algorithm, and a Bayesian optimization strategy is used to search for the optimal parameter combination of the LightGBM algorithm.
[0089] The beneficial effects of the above technical solution are as follows: This solution makes full use of big data mining and machine learning technology to identify key factors affecting scheduling decisions from massive random simulated runoff and corresponding optimized scheduling results. While ensuring the effectiveness of machine learning in extracting reservoir group scheduling rules, it greatly reduces the cost of machine learning training and provides reliable technical support for the coordinated operation and decision-making of multi-reservoir systems.
[0090] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0091] First, this invention proposes a method system for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff, realizing a complete technical path of "two-dimensional spatiotemporal correlation analysis of runoff - stochastic simulation of runoff in multiple reservoirs - optimized scheduling of reservoir groups - data mining of key decision factors - machine learning scheduling rule extraction".
[0092] Secondly, unlike traditional scheduling rule extraction methods, this invention makes full use of big data mining and machine learning technologies, which can identify key factors affecting scheduling decisions from a large number of historical and optimized scheduling results. While ensuring the effectiveness of machine learning in extracting reservoir group scheduling rules, it greatly reduces the cost of machine learning training, improves the scientificity and intelligence of reservoir group scheduling rule extraction, and provides reliable technical support for the joint operation and decision-making of multi-reservoir systems.
[0093] Third, unlike most existing stochastic runoff simulation studies that focus on temporal correlation, this invention considers the two-dimensional correlation characteristics of runoff in time and space, and establishes a stochastic simulation method for multi-reservoir runoff. It uses the Monte Carlo inverse transform sampling method to randomly simulate and generate massive amounts of multi-reservoir runoff data with temporal and spatial correlation, effectively making up for the limited time length of historical runoff monitoring data. It is also applicable to different spatial topological connections such as series, parallel, and mixed connections, providing a more scientific and reliable data foundation for the scheduling of reservoir groups with hydraulic and electrical connections.
[0094] Fourth, this invention proposes a multi-scenario runoff construction method that combines conventional and extreme scenarios. Based on conventional scenario multi-reservoir long-sequence simulated runoff, through extreme runoff identification and spatiotemporal clustering analysis, it further filters out different types of extreme scenario multi-reservoir long-sequence simulated runoff considering spatiotemporal two-dimensional correlation, ensuring the applicability of reservoir group scheduling rules under both conventional and extreme conditions.
[0095] In the process of constructing the reservoir group scheduling rule model, this scheme can be adapted to the continuous accumulation of hydrological data and scheduling practice data. Practice has yielded a wealth of scheduling experience and wisdom. By applying data mining and machine learning to this experience and wisdom, and combining the optimized scheduling results with actual scheduling results, a scheduling rule extraction model can be established. This model can better adapt to new changes and new needs, thereby better supporting reservoir group scheduling decisions. Attached Figure Description
[0096] Figure 1 The flowchart shows a method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff.
[0097] Figure 2 This is a schematic diagram illustrating the spatial and temporal correlations of runoff in a reservoir group.
[0098] Figure 3 Comparison of fitting results for different marginal distribution functions: (a) Comparison of fitting results for the marginal distribution function of runoff in reservoir 1 during a certain period; (b) Comparison of fitting results for the marginal distribution function of runoff in reservoir 2 during a certain period.
[0099] Figure 4The fitting results of the joint probability distribution function of runoff Copula for the same reservoir at different time periods and for the same time period at different reservoirs are as follows: (a) The fitting results of the joint probability distribution function of runoff Frank Copula for the same reservoir at time periods t and t+1; (b) The fitting results of the joint probability distribution function of runoff Frank Copula for the same reservoir at time period t for the same reservoir 1 and 2.
[0100] Figure 5 Comparison of root mean square error (RMSE) of fitting results for different Copula joint distribution functions: (a) Comparison of RMSE for time period t and t+1 of reservoir 1, (b) Comparison of RMSE for time period t and t+1 of reservoir 2, (c) Comparison of RMSE for time period t of reservoir 1 and reservoir 2.
[0101] Figure 6 The results of the stochastic simulation of runoff considering the two-dimensional spatiotemporal correlation are as follows: (a) is the stochastic simulation result of runoff for reservoir 1; (b) is the stochastic simulation result of runoff for reservoir 2.
[0102] Figure 7 A flowchart for calculating runoff in a long-sequence simulation of multiple reservoirs under extreme scenarios that take into account spatiotemporal two-dimensional correlation.
[0103] Figure 8 This is a comparison chart of the grey relational degrees of different candidate scheduling decision factors based on data mining using grey relational analysis.
[0104] Figure 9 This diagram illustrates the input and output of the LightGBM machine learning method for extracting scheduling rules.
[0105] Figure 10 The diagram shows a comparison of the results of three scheduling methods: scheduling based on traditional scheduling graphs, scheduling based on rules extracted by LightGBM machine learning, and scheduling directly optimized by the GA algorithm. (a) Comparison of the root mean square error (RMSE) of the outflow results of scheduling based on traditional scheduling graphs and scheduling based on rules extracted by LightGBM machine learning with scheduling directly optimized by the GA algorithm. (b) Comparison of the power generation results of the three scheduling methods: scheduling based on traditional scheduling graphs, scheduling based on rules extracted by LightGBM machine learning, and scheduling directly optimized by the GA algorithm. Detailed Implementation
[0106] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0107] refer to Figure 1 , Figure 1 A flowchart illustrating a method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff is shown, as follows: Figure 1 As shown, the method S includes steps S1 to S5.
[0108] refer to Figure 2 , Figure 2 Taking a series of reservoirs as an example, the spatial and temporal correlations of runoff in a reservoir group are illustrated. Figure 2 In the spatial relationship on the left, if the runoff range is discretized into M intervals, then the reservoir runoff at time i t has the following possibilities: , … The reservoir runoff at time i+1 t has the following possibilities: , … The following combinations of the two are possible. and , and … and , and , and … and … and , and … and ; Figure 2 In the time relationship on the right, the runoff range is also discretized into M intervals. Then, the runoff of the reservoir in time period t at time i has the following possibilities: , … The reservoir runoff at time t+1 has the following possibilities: , … The following combinations of the two are possible. and , and … and , and , and … and … and , and … and .pass Figure 2 It can be seen that the inflow of the reservoir group is correlated and interconnected in time and space, and has multiple combinations. Each combination corresponds to a different probability. If the reservoir group scheduling rules are established using limited hydrological monitoring data, it is impossible to exhaust all combinations, which reduces the applicability of the reservoir group scheduling rules.
[0109] In step S1, a group of reservoirs with hydraulic and electrical connections is taken as the research object. The runoff margin distribution function of each reservoir and each time period is screened. Based on the runoff margin distribution function, the runoff Copula joint distribution function of the same reservoir at different time periods and the same time period at different reservoirs is screened.
[0110] In one embodiment of the present invention, the detailed implementation method of step S1 includes:
[0111] S11. Select multiple probability distribution functions, including at least the Normal function, Gamma function, Log-normal function, GEV function, and Weibull function, and calculate the runoff marginal distribution function for each reservoir and each time period. Their specific expressions are as follows:
[0112] A1. The Normal function is:
[0113]
[0114] Where x is a random variable; μ is the mean of the normal distribution; σ is the standard deviation of the normal distribution; σ 2 The variance is the normal distribution.
[0115] A2. The Gamma function is:
[0116]
[0117] Where α is the shape parameter, which determines the thickness of the tail of the distribution; β is the scale parameter, which controls the width of the distribution; Γ(α) is the Gamma function, used to normalize the probability density function;
[0118] A3. The Log-normal function is:
[0119]
[0120] in, The mean of a log-normal distribution; The standard deviation of the log-normal distribution; 2 The variance of the log-normal distribution;
[0121] A4. The GEV function is:
[0122]
[0123] Where y is the probability density value of the function; The shape parameter determines the tail thickness of the distribution; This is a scale parameter that controls the width of the distribution; This is a location parameter used to adjust the offset of the distribution;
[0124] A5. The Weibull function is:
[0125]
[0126] Where ω is a shape parameter that determines the shape of the distribution; λ is a scale parameter that determines the width and location of the distribution.
[0127] refer to Figure 3 , Figure 3 A comparison chart of fitting results for different marginal distribution functions is shown, in which... Figure 3 (a) is a comparison of the fitting results of the runoff margin distribution function of Reservoir 1 during a certain period. Figure 3 (b) is a comparison of the fitting results of the runoff edge distribution function of Reservoir 2 during a certain period.
[0128] S12. The parameters of the runoff marginal distribution function are calculated using the maximum likelihood method, and the root mean square error is used as the evaluation index to verify the results of various distribution functions and select the most suitable runoff marginal distribution function.
[0129] The specific expression for the root mean square error (RMSE) is as follows:
[0130]
[0131] in, The number of samples; For the first One measured value; It is the first The simulation values are: RMSE, which ranges from [0, +∞). The closer the RMSE is to 0, the better the simulation effect of the model.
[0132] S13. Based on the runoff edge distribution function, select multiple runoff Copula joint distribution functions, and calculate the runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs respectively;
[0133] The Copula function, used to describe the correlation of multiple variables, is a function that connects the joint distribution function with its respective marginal distribution functions; it can also be called a link function. Taking the two-dimensional joint distribution function as an example, it can be expressed as:
[0134]
[0135] Where C is a Copula function; For the parameters of the Copula function; u= v= Let X and Y represent the marginal distribution functions of random variables X and Y, respectively.
[0136] The Copula joint distribution function includes at least the Clayton Copula function, the Frank Copula function, and the Gumbel Copula function, and their expressions are as follows:
[0137] A1. The Clayton Copula function is:
[0138]
[0139] A2. The Frank Copula function is:
[0140]
[0141] A3. The Gumbel Copula function is:
[0142]
[0143] Among them, C and These are the Copula function and its parameters, respectively.
[0144] refer to Figure 4 , Figure 4 The fitting results of the joint probability distribution function of runoff Copula for the same reservoir at different time periods and for the same time period at different reservoirs are shown. Figure 4 (a) shows the fitting results of the Frank Copula joint probability distribution function of runoff in the same reservoir at time periods t and t+1. Figure 4 (b) shows the Frank Copula joint probability distribution function fitting results for the runoff of reservoirs 1 and 2 during the same time period t.
[0145] S14. The parameters of the runoff Copula joint distribution function are calculated using the maximum likelihood method. The root mean square error is used as the evaluation index to verify the results of various Copula joint distribution functions and select the most suitable runoff Copula joint distribution function for the same reservoir at different time periods and for the same time period at different reservoirs.
[0146] refer to Figure 5 , Figure 5 The graph shows a comparison of the root mean square error (RMSE) of fitting results for different Copula joint distribution functions. Figure 5 (a) is a comparison chart of RMSE for reservoir 1 at time t and t+1. Figure 5 (b) is a comparison chart of RMSE for reservoir time periods t and t+1. Figure 5 (c) is a comparison chart of RMSE for time period t for reservoir 1 and reservoir 2.
[0147] Based on the optimized runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs. Let any reservoir i∈[1,I] and any time period t∈[1,T]. When establishing the joint distribution of runoff for the same reservoir at different (adjacent) time periods, u and v are the marginal distribution functions for any reservoir i at time period t and time period t+1, respectively. When establishing the joint distribution of runoff for different (adjacent) reservoirs at the same time period, u and v are the marginal distribution functions for any time period t for reservoir i and reservoir i+1, respectively.
[0148] Given the runoff of the reservoir during time period t in time i, find the expression for the runoff of the reservoir during time period t+1 in time i. for:
[0149]
[0150] in, Let be the joint Copula distribution function of runoff for any reservoir i at time t and time t+1; and Let be the runoff of any reservoir i at time t and time t+1, respectively.
[0151] Given the runoff of reservoir i in time period t, find the expression for the runoff of reservoir i+1 in time period t. for:
[0152]
[0153] in, Let be the joint Copula distribution function of runoff from reservoir i and reservoir i+1 at any time interval t; and Let t represent the runoff of reservoir i and reservoir i+1 at any given time period t.
[0154] S2. Considering the two-dimensional spatiotemporal characteristics of runoff, using the selected runoff Copula joint distribution function and combining the inverse transformation sampling of the Monte Carlo method, multiple sets of runoff sequence data containing all reservoirs and all time periods are generated through multiple random simulations.
[0155] The main steps for generating multiple sets of conventional scenario runoff sequence data covering all reservoirs and all time periods include steps S21 to S23:
[0156] S21. Randomly simulate the runoff of Reservoir 1 at different (adjacent) time periods, then:
[0157] ① Use a linear congruential random number generator to generate random numbers a1∈(0,1);
[0158] ② Use the inverse function of the cumulative distribution function in the marginal distribution to map random numbers to random variables of the target distribution, i.e., solve for... The runoff x of reservoir 1 during time period 1 can be obtained. 1,1 ;
[0159] ③ Generate random numbers a2, a3, ..., a T Given (0,1), we can obtain the following system of equations:
[0160]
[0161] Solving the above system of equations will yield the runoff (x) of the reservoir during time period 2 to T. 1,2 , x 1,3 , …, x 1,T-1 , x 1,T );
[0162] S22. Randomly simulate the runoff of Reservoir 2 at different (adjacent) time periods, then:
[0163] ① Use a linear congruential random number generator to generate random numbers b1∈(0,1);
[0164] ② Order Given the runoff x of reservoir 1 during time period 1. 1,1 Solving this equation will yield the runoff x during time period 1 of reservoir 2. 2,1 ;
[0165] ③ Generate random numbers b2, b3, ..., b T Given (0,1), we can obtain the following system of equations:
[0166]
[0167] Solving the above system of equations will yield the runoff (x) of the reservoir during time period 2~T. 2,2, x 2,3 , …, x 2,T-1 , x 2,T );
[0168] S23. Following S22, perform runoff simulations for reservoirs 3 through I in sequence.
[0169] By completing the above steps, we can obtain conventional scenario multi-reservoir runoff sequence data that considers spatiotemporal two-dimensional correlation features.
[0170] refer to Figure 6 , Figure 6 The results of stochastic runoff simulation considering spatiotemporal two-dimensional correlations are shown, where Figure 6 (a) shows the stochastic simulation results of runoff in Reservoir 1. Figure 6 (b) shows the stochastic simulation results of runoff in Reservoir 2.
[0171] To generate multiple sets of extreme scenario runoff sequence data covering all reservoirs and all time periods, the main steps further include step S24:
[0172] S24. Based on the generated multiple sets of long-sequence simulated runoff from multiple reservoirs under normal scenarios, further screening is conducted through extreme runoff identification and spatiotemporal clustering analysis to obtain different types of long-sequence simulated runoff from multiple reservoirs under extreme scenarios that consider spatiotemporal two-dimensional correlations. The specific steps are as follows:
[0173] A1. Extreme runoff identification markers based on frequency analysis
[0174] Frequency statistical analysis is performed on historical runoff data for all reservoirs and all time periods to define extreme runoff thresholds for any reservoir and any time period (for example, for reservoir i in time period t, runoff with a frequency P ≤ 10% is defined as the extreme abundant runoff threshold, and runoff with a frequency P ≥ 90% is defined as the extreme dry runoff threshold). Then, using these thresholds as the criterion, multiple sets of long-sequence simulated runoff data from multiple reservoirs under conventional scenarios are identified and labeled, allowing for the selection of long-sequence simulated runoff data from multiple reservoirs under extreme scenarios that considers spatiotemporal two-dimensional correlation.
[0175] A2. Extreme runoff clustering analysis based on temporal and spatial dimensions
[0176] After selecting and obtaining long-sequence simulated runoff of multiple reservoirs under extreme scenarios that consider the two-dimensional correlation of time and space, cluster analysis of the extreme scenarios can be further performed from both temporal and spatial perspectives.
[0177] B21. Time-based cluster analysis:
[0178] ①Single-period extreme scenario: For a set of runoff sequences that include all reservoirs and all time periods, if the runoff in any time period reaches the extreme threshold, it is defined as "single-period extreme abundance" or "single-period extreme scarcity".
[0179] ② Multi-period extreme scenario: For a set of runoff sequences that include all reservoirs and all time periods, if the runoff in more than one time period reaches the extreme threshold, it is defined as "multi-period extreme abundance", "multi-period extreme scarcity" or "multi-period abundance-scarcity shift".
[0180] B22. Spatial Cluster Analysis:
[0181] ① Extreme scenario for a single reservoir: For a set of runoff sequences that include all reservoirs and all time periods, if the runoff of any reservoir reaches the extreme threshold, it is defined as "extreme abundance of a single reservoir", "extreme drought of a single reservoir" or "transition between abundance and drought of a single reservoir".
[0182] ②Multiple reservoir extreme scenario: For a set of runoff sequences that include all reservoirs and all time periods, if the runoff of more than one reservoir reaches the extreme threshold, it is defined as "multiple reservoir extreme abundance", "multiple reservoir extreme drought" or "multiple reservoir abundance-dryness shift".
[0183] The following example illustrates this with a runoff sequence involving two reservoirs over 12 months:
[0184] ①If only the July runoff of Reservoir 2 reaches the extreme abundance threshold, then the time clustering of this sequence is labeled as "single-period extreme abundance" and the spatial clustering is labeled as "single-reservoir extreme abundance";
[0185] ②If only the April runoff of Reservoir 1 reaches the extreme low threshold, then the time clustering of this sequence is labeled as "single-period extreme low" and the spatial clustering is labeled as "single-reservoir extreme low".
[0186] ③ If only the runoff of Reservoir 2 reaches the extreme abundance threshold in July and August, then the time clustering of this sequence is labeled as "multi-period extreme abundance" and the spatial clustering is labeled as "single reservoir extreme abundance".
[0187] ④ If only the runoff of Reservoir 1 reaches the extreme low threshold in April and May, then the time clustering of this sequence is labeled as "multi-period extreme low" and the spatial clustering is labeled as "single reservoir extreme low".
[0188] ⑤ If only the runoff of Reservoir 1 reaches the extreme low threshold in June and the runoff reaches the extreme high threshold in July, then the time clustering of this sequence is labeled as "multi-period high-low transition" and the spatial clustering is labeled as "single reservoir high-low transition".
[0189] ⑥ If the runoff of Reservoir 1 and Reservoir 2 only reaches the extreme abundance threshold in July, then the time clustering of this sequence is labeled as "single-period extreme abundance" and the spatial clustering is labeled as "multiple reservoirs extreme abundance".
[0190] ⑦ If the runoff of Reservoir 1 and Reservoir 2 only reaches the extreme low threshold in April, then the time clustering of this sequence is labeled as "single-period extreme low" and the spatial clustering is labeled as "multiple reservoirs extreme low".
[0191] ⑧ If the runoff of Reservoir 1 and Reservoir 2 reaches the extreme abundance threshold in June, and the runoff of Reservoir 1 and Reservoir 2 reaches the extreme abundance threshold in July, then the time clustering of this sequence is labeled as "multi-period extreme abundance" and the spatial clustering is labeled as "multi-reservoir extreme abundance".
[0192] ⑨ If the runoff of Reservoir 1 and Reservoir 2 reaches the extreme low threshold in March, and the runoff of Reservoir 1 and Reservoir 2 reaches the extreme low threshold in April, then the time clustering of this sequence is labeled as "multi-period extreme low" and the spatial clustering is labeled as "multi-reservoir extreme low".
[0193] ⑩ If the runoff of Reservoir 1 and Reservoir 2 reaches the extreme low threshold in June, and the runoff of Reservoir 1 reaches the extreme high threshold in July, then the time clustering of this sequence is labeled as "multi-period high-low transition" and the spatial clustering is labeled as "multi-reservoir high-low transition".
[0194] ⑪ If the runoff of Reservoir 1 reaches the extreme low threshold in June and the runoff of Reservoir 2 reaches the extreme high threshold in July, then the time clustering of this sequence is labeled as "multi-period high-low transition" and the spatial clustering is labeled as "multi-reservoir high-low transition".
[0195] refer to Figure 7 , Figure 7 A flowchart illustrating the calculation process for long-sequence simulated runoff from multiple reservoirs under extreme scenarios, considering spatiotemporal two-dimensional correlations, is presented. By combining long-sequence simulated runoff from multiple reservoirs under conventional scenarios and employing the aforementioned extreme runoff identification and spatiotemporal clustering analysis, different types of long-sequence simulated runoff from multiple reservoirs under extreme scenarios, considering spatiotemporal two-dimensional correlations, can be obtained.
[0196] In step S3, the objective function and constraints of the reservoir group optimization are determined, and the reservoir group optimization scheduling model is established. Multiple sets of runoff sequence data from multiple random simulations are used as model inputs, and the optimization algorithm is used to solve the model to obtain the corresponding multiple sets of reservoir group optimization scheduling results, including all reservoirs, the end-of-period reservoir capacity of all time periods, the average outflow of time periods, the average power generation flow of time periods, the average water discharge of time periods, and the average power output of time periods.
[0197] For ease of understanding, this invention uses maximizing power generation as the objective function for optimizing the reservoir group, and its expression is as follows:
[0198] ;
[0199] ;
[0200] ;
[0201] ;
[0202] ;
[0203] ;
[0204] in, The objective is to maximize the power generation of a group of reservoirs with both hydroelectric and electrical connections; i is the reservoir index. I represents the total number of reservoirs in the reservoir group; t represents the time period index. T represents the total number of time periods; For the power station of reservoir i during the time period The power generation capacity; The time period is long; The density of water; It is the acceleration due to gravity; Let be the power generation efficiency of the power station in reservoir i; , and These represent the outflow from reservoir i, the power generation flow from the power station, and the wastewater discharge during time period t, respectively. , , and These represent the average head, average water level in front of the dam, average water level behind the dam, and average head loss of the power station at reservoir i during time period t. Let i be the average reservoir capacity of reservoir i during time period t; and These represent the reservoir capacity of reservoir i at the beginning and end of time period t, respectively. , , , and and , , , and All coefficients are obtained using polynomial fitting;
[0205] The constraints include the water balance equation, initial / terminal reservoir capacity constraints, reservoir capacity constraints, outflow constraints, power generation flow constraints, and hydropower station output constraints.
[0206] The expression for the water balance equation is:
[0207] ;
[0208] The expression for the initial / termination storage capacity constraint is:
[0209] ; ;
[0210] The expression for the reservoir capacity constraint is:
[0211] ;
[0212] The expression for the outflow constraint is:
[0213] ;
[0214] The expression for the power generation flow constraint is:
[0215] ;
[0216] The expression for the power output constraint of the hydropower station is:
[0217] ;
[0218] in, For reservoir During the period Inbound traffic; For reservoir During the period Evaporation rate; For reservoir Storage capacity at the beginning of the scheduling period; For reservoir The target storage capacity at the end of the scheduling period; and Reservoirs During the period Minimum and maximum storage capacity; and Reservoirs During the period The minimum and maximum discharge flow rates; and Reservoirs The power station during the period The minimum and maximum power generation flow rates; and Reservoirs The power station during the period The minimum and maximum power generation capacity;
[0219] In one embodiment of the present invention, the optimization algorithm is a genetic algorithm, and its detailed implementation method includes:
[0220] A1. Initialize the parameters of the genetic algorithm, using multiple sets of runoff sequence data from all reservoirs and all time periods from multiple random simulations in step S2 as model inputs (normal scenario or extreme scenario).
[0221] A2. Using the total discharge flow of all reservoirs, the average discharge flow over all time periods, or the reservoir capacity at the end of a time period as decision variables, randomly generate N different decision variables as the initial population. ;
[0222] A3. Calculate the fitness of each individual in the population based on the objective function of the reservoir group optimal scheduling model. The expression for the objective function is:
[0223]
[0224] in, The objective is to maximize the power generation of a group of reservoirs with both hydroelectric and electrical connections; i is the reservoir index. t is the time period index. ; For the power station of reservoir i during the time period The power generation capacity; The time period is long.
[0225] A4. Use the tournament selection algorithm to select individuals from the population to generate a population of size N / 2. ;
[0226] A5. According to the preset crossover probability, select from the population Individuals are selected for crossover to generate offspring, and these offspring replace the original population. The parent individual in the middle;
[0227] A6. Adjust the population according to the preset mutation probability. Individuals undergo mutation operations, and the mutated individuals replace the original individuals in the next generation of the population. Chromosomes that have not undergone mutation directly enter the next generation of the population. ;
[0228] A7. Calculate the population based on the objective function of the reservoir group optimal scheduling model. The fitness value of each individual is calculated, and the current generation number Gen is incremented by 1;
[0229] A8. Determine whether the number of generations Gen exceeds the set maximum number of generations or whether all fitness values meet the set threshold. If yes, proceed to step A9; otherwise, return to step A3.
[0230] A9. Select the individual with the best fitness in the population as the optimal decision variable, and output the operating parameters corresponding to the optimal decision variable, such as the reservoir capacity at the end of the time period, the average discharge flow, the average power generation flow, the average water abandonment flow, and the average power output, for all reservoirs and all time periods. This is the optimized scheduling result.
[0231] In step S4, based on multiple sets of runoff sequence data from multiple random simulations and their corresponding multiple sets of reservoir group optimization scheduling calculation results, key decision factors are determined using data mining methods, and a reservoir group scheduling rule model is established using machine learning methods.
[0232] In one embodiment of the present invention, data mining is performed using grey relational analysis (GRA), and the specific steps include:
[0233] A1. Constructing a multi-scheme sample set for data mining
[0234] Using multiple sets of runoff data (normal or extreme scenarios) from the above-mentioned random simulations and the corresponding results of reservoir group optimization scheduling calculations, a multi-scheme sample set for data mining is constructed. Each scheme sample should include all reservoirs, initial reservoir capacity for all time periods, final reservoir capacity for each time period, average inflow for each time period, average outflow for each time period, average power generation for each time period, average water discharge for each time period, and average power output for each time period.
[0235] A2. Use data mining methods to identify key decision factors.
[0236] B1. Calculate the grey relational degree of each decision factor.
[0237] The average outflow or end-of-period reservoir capacity of all reservoirs is selected as the decision variable. The position of the period in the year, as well as the inflow, initial reservoir capacity, and average outflow of several previous and current periods of all reservoirs, are selected as candidate decision factors. The grey relational degree between the candidate decision factors and the decision variables is calculated one by one. The specific expression is as follows:
[0238]
[0239] in, The grey relational degree of candidate decision factor j; The grey relational coefficient of candidate decision factor j in record n; [0,N], where N is the total number of records (e.g., if 10,000 runoff schemes are randomly generated, and each scheme covers 12 months, then N=120,000). The value of the decision variable at record n; Let j be the value of the candidate decision factor j at record n; The resolution coefficient is the standard value. Set it to 0.5. Wherein, the grey relational degree... The value range is [0,1]. The closer the value is to 1, the stronger the correlation between the two variables; the closer the value is to 0, the weaker the correlation between the two variables.
[0240] B2. Set the gray relational degree threshold (e.g., 0.75), select the value that satisfies the condition. The factors are used as key decision factors.
[0241] refer to Figure 8 , Figure 8 This paper presents a comparison of the grey relational degrees of different candidate scheduling decision factors based on grey relational analysis. The decision variable is the average outflow rate for the current period, while the candidate decision factors include the month of the period, the inflow rate for the current period and several previous periods, the initial storage capacity for the current period and several previous periods, and the average outflow rate for several previous periods. Due to the candidate decision factors... and satisfy Therefore, it was selected as a key decision factor.
[0242] Through the above steps, the key decision factors for reservoir group scheduling can be screened out, which will enable the subsequent use of machine learning methods to establish a reservoir group scheduling rule model, ensuring the decision support effect of the scheduling rule model while reducing the dimensionality of decision factors.
[0243] In one embodiment of the present invention, machine learning is performed using the LightGBM algorithm, and the specific steps include:
[0244] A1. Constructing and partitioning multi-scheme sample sets for machine learning models
[0245] Using the decision variables determined in the above steps and the key decision factors screened out by grey relational analysis, a multi-scheme sample set for the machine learning model is further constructed. The sample set is divided into a training set and a test set according to a certain ratio (such as 7:3 or 8:2). The training set is used for training the machine learning model and optimizing parameters, while the test set is used for model accuracy verification and generalization ability evaluation.
[0246] A2. Establishment of a reservoir group scheduling rule model based on machine learning methods
[0247] Based on the training set of the multi-scheme sample set constructed above, the LightGBM machine learning method is used to extract the reservoir group scheduling rules. Compared with the traditional layer-by-layer growth strategy of decision trees, the LightGBM algorithm adopts a leaf-wise growth strategy with depth constraints. This strategy finds the leaf with the largest split gain from all current leaves and splits it, repeating this process. In this way, the LightGBM algorithm can gradually optimize the tree structure, accelerate the learning process, and improve the model accuracy. The specific expression for the split gain (Gain) of the node to be split is as follows:
[0248]
[0249] Wherein, Node is the node to be split; and These represent the left and right child nodes generated after a node (Node) splits; Let be the loss value for node Node. and This represents the loss values for the left and right subtrees after the split. If Gain(Node) > 0, it indicates that the split effectively reduces the overall loss, and the split operation should be performed; otherwise, the split should not be performed.
[0250] The LightGBM machine learning method extracts reservoir group scheduling rules. Its performance is highly dependent on the selection of multiple parameters, which need to be optimized during model training. These parameters mainly include: maximum tree depth (max_depth), number of leaf nodes (num_leaves), learning rate (learning_rate), number of iterations (n_estimators), feature sampling ratio (feature_fraction), data sampling ratio (bagging_fraction), minimum number of samples per leaf node (min_data_in_leaf), regularization parameters (lambda_l1 / lambda_l2), and minimum split gain (min_split_gain). This embodiment uses a Bayesian optimization strategy to search the parameter space, using the root mean square error (RMSE) as the evaluation metric to obtain the optimal parameter combination.
[0251] refer to Figure 9 , Figure 9 The diagram illustrates the input and output of the LightGBM machine learning method for extracting scheduling rules. The target variable is the average outflow rate for the current period, and the feature variables include the month of the period, the inflow rate for the current period and several previous periods, the initial storage capacity for the current period and several previous periods, and the average outflow rate for several previous periods.
[0252] A3. Verification of Reservoir Group Scheduling Rule Model Based on Machine Learning Methods
[0253] This invention validates the model using a test set, employing root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The effectiveness of reservoir group scheduling rules based on the LightGBM machine learning method is evaluated using indicators such as [list of indicators]. Furthermore, to ensure the feasibility of the rule extraction model's predictions, this invention further substitutes the predicted reservoir group's time-period decision variables (time-period average outflow or time-end reservoir capacity) into the corresponding reservoir operation constraints for verification, quantifying the constraint compliance to determine whether the prediction results can be used for scheduling operations. For any reservoir in the reservoir group, the feasibility evaluation indicators are as follows:
[0254] Storage capacity constraint over-limit rate (VCR):
[0255]
[0256] Outflow over-limit rate (OCR):
[0257]
[0258] Typically, VCR ≤ 5% and OCR ≤ 3% are required.
[0259] If it is verified that all reservoirs in the reservoir group meet the above evaluation index requirements, then the reservoir group scheduling rule model based on the LightGBM machine learning method can be considered to be effective and feasible.
[0260] S5. Using the state variables of the current and several preceding / following time periods of the determined key decision factors as input to the reservoir group scheduling rule model, the scheduling operation results of the reservoir group are predicted. The state variables include the position of the time period in a year and the measured and predicted inflow of all reservoirs in the current and several preceding / following time periods, the measured initial reservoir capacity in the current and several preceding time periods, and the measured average outflow in the several preceding time periods. The scheduling operation results include the final reservoir capacity of all reservoirs in the current and several preceding time periods, the average outflow, the average power generation flow, the average water discharge flow, and the average power output.
[0261] In one embodiment of the present invention, during scheduling practice, the state quantities of each key decision factor at the current time and several preceding and following time periods (including the position of the time period in a year, the measured and predicted inflow of all reservoirs at the current time and several preceding and following time periods, the measured initial reservoir capacity at the current time and several preceding time periods, the measured average outflow at several preceding time periods, etc.) are used as inputs to a reservoir group scheduling rule model based on the LightGBM machine learning method. Through model calculation, the reservoir group scheduling operation results (including the final reservoir capacity, average outflow, average power generation flow, average water discharge, and average power output of all reservoirs at the current time and several preceding time periods) can be obtained, thereby providing decision support for reservoir group scheduling. In addition, during scheduling practice, the rule extraction model can be used to perform rolling predictions on a time-by-time basis, dynamically responding to changes in inflow and engineering operating conditions, providing scientific prediction results, and better supporting reservoir group scheduling decisions.
[0262] refer to Figure 10 , Figure 10The diagram shows a comparison of the results of three scheduling methods: scheduling based on traditional scheduling graphs, scheduling based on LightGBM machine learning extraction rules, and scheduling directly optimized using the GA algorithm. Specifically, the root mean square error (RMSE) of the outflow results for scheduling based on traditional scheduling graphs and scheduling based on LightGBM machine learning extraction rules is compared with that of scheduling directly optimized using the GA algorithm. The results show that scheduling based on LightGBM machine learning extraction rules is superior to scheduling based on traditional scheduling graphs. A comparison of the power generation results for these three methods shows that scheduling based on LightGBM machine learning extraction rules is closer to scheduling directly optimized using the GA algorithm than scheduling based on traditional scheduling graphs.
Claims
1. A method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff, characterized in that, Including the following steps: S1. Taking a group of reservoirs with hydraulic and electrical connections as the research object, we screen the runoff margin distribution function for each reservoir and each time period, and screen the runoff Copula joint distribution function for the same reservoir and different time periods, as well as for different reservoirs in the same time period, based on the runoff margin distribution function. S2. Considering the two-dimensional spatiotemporal characteristics of runoff, using the selected runoff Copula joint distribution function and combining the inverse transformation sampling of the Monte Carlo method, multiple sets of runoff sequence data containing all reservoirs and all time periods are generated through multiple random simulations. S3. Determine the objective function and constraints for reservoir group optimization, establish a reservoir group optimization scheduling model, use multiple sets of runoff sequence data from multiple random simulations as model input, use optimization algorithms to solve the model, and obtain the corresponding multiple sets of reservoir group optimization scheduling results, including all reservoirs, the end-of-period reservoir capacity for all time periods, the average outflow for each time period, the average power generation flow for each time period, the average water discharge for each time period, and the average power output for each time period. S4. Based on multiple sets of runoff sequence data from multiple random simulations and their corresponding multiple sets of reservoir group optimization scheduling calculation results, data mining methods are used to determine key decision factors, and machine learning methods are used to establish a reservoir group scheduling rule model. S5. Using the state variables of the current and several preceding and following time periods of the determined key decision factors as input to the reservoir group scheduling rule model, the reservoir group scheduling operation results are predicted; the state variables include the position of the time period in a year and the measured and predicted inflow of all reservoirs in the current and several preceding and following time periods, the measured initial reservoir capacity in the current and several preceding time periods, and the measured average outflow in the several preceding time periods. Step S1 further includes: S11. Select multiple probability distribution functions, including at least the Normal function, Gamma function, Log-normal function, GEV function and Weibull function, and calculate the runoff marginal distribution function for each reservoir and each time period respectively; S12. The parameters of the runoff marginal distribution function are calculated using the maximum likelihood method, and the root mean square error is used as the evaluation index to test the results of various distribution functions and select the most suitable runoff marginal distribution function. S13. Based on the runoff edge distribution function, select multiple runoff Copula joint distribution functions, including at least the Clayton Copula function, the Frank Copula function, and the Gumbel Copula function, and calculate the runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs respectively. S14. The parameters of the runoff Copula joint distribution function are calculated using the maximum likelihood method, and the root mean square error is used as the evaluation index to verify the results of various runoff Copula joint distribution functions, and the most suitable runoff Copula joint distribution functions for the same reservoir at different time periods and for the same time period at different reservoirs are selected. Given the runoff of the reservoir during time period t in time i, find the expression for the runoff of the reservoir during time period t+1 in time i. for: in, Let be the joint Copula distribution function of runoff for any reservoir i at time t and time t+1; and Let be the runoff of any reservoir i at time t and time t+1, respectively; Given the runoff of reservoir i in time period t, find the expression for the runoff of reservoir i+1 in time period t. for: in, Let be the joint Copula distribution function of runoff from reservoir i and reservoir i+1 at any time interval t; and Let t represent the runoff of reservoir i and reservoir i+1 at any given time period; The objective function for optimizing the reservoir group is to maximize power generation, and its expression is: ; ; ; ; ; ; in, The objective is to maximize the power generation of a group of reservoirs with both hydroelectric and electrical connections; i is the reservoir index. I represents the total number of reservoirs in the reservoir group; t represents the time period index. T represents the total number of time periods; For the power station of reservoir i during the time period The power generation capacity; For a long period of time; The density of water; It is the acceleration due to gravity; Let be the power generation efficiency of the power station in reservoir i; , and These represent the outflow from reservoir i, the power generation flow from the power station, and the wastewater discharge during time period t, respectively. , , and These represent the average head, average water level in front of the dam, average water level behind the dam, and average head loss of the power station at reservoir i during time period t. Let i be the average reservoir capacity of reservoir i during time period t; and These represent the reservoir capacity of reservoir i at the beginning and end of time period t, respectively. , , , and and , , , and All coefficients are obtained using polynomial fitting; The constraints include water balance equations, initial / terminal reservoir capacity constraints, reservoir capacity constraints, outflow constraints, power generation flow constraints, and hydropower station output constraints. The expression for the water balance equation is: ; The expression for the initial / termination storage capacity constraint is: ; ; The expression for the reservoir capacity constraint is: ; The expression for the outflow constraint is: ; The expression for the power generation flow constraint is: ; The expression for the power output constraint of the hydropower station is: ; in, For reservoir During the period Inbound traffic; For reservoir During the period Evaporation rate; For reservoir Storage capacity at the beginning of the scheduling period; For reservoir The target storage capacity at the end of the scheduling period; and Reservoirs During the period Minimum and maximum storage capacity; and Reservoirs During the period The minimum and maximum discharge flow rates; and Reservoirs The power station during the period The minimum and maximum power generation flow rates; and Reservoirs The power station during the period The minimum and maximum power generation capacity.
2. The method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff as described in claim 1, characterized in that, The detailed implementation method for generating multiple sets of runoff sequence data containing all reservoirs and all time periods in step S2 includes: S21. To randomly simulate the runoff of Reservoir 1 in adjacent time periods, a linear congruent random number generator is used to generate random numbers a1∈(0,1); the inverse function of the cumulative distribution function in the marginal distribution is used to map the random numbers to random variables of the target distribution, and the solution is obtained. The runoff x of reservoir 1 during time period 1 can be obtained. 1,1 ; Generate random numbers a2, a3, ..., a T Given (0,1), we can obtain the following system of equations: Solving the above system of equations will yield the runoff x of the reservoir during time period 2 to T. 1,2 , x 1,3 , ..., x 1,T-1 , x 1,T ; S22. To randomly simulate the runoff of Reservoir 2 in adjacent time periods, a linear congruent random number generator is used to generate random numbers b1∈(0,1); let Given the runoff x of reservoir 1 during time period 1. 1,1 Solving this equation will yield the runoff x during time period 1 of reservoir 2. 2,1 ; Generate random numbers b2, b3, ..., b T Given (0,1), we can obtain the following system of equations: Solving the above system of equations will yield the runoff x of the reservoir during time period 2 to T. 2,2 , x 2,3 , ..., x 2,T-1 , x 2,T ; S23. Following S22, perform runoff simulations for reservoirs 3 through I in sequence.
3. The method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff as described in claim 2, characterized in that, Step S2 also includes generating multiple sets of extreme scenario runoff sequence data containing all reservoirs and all time periods, the implementation method of which further includes: Frequency statistical analysis was performed on historical runoff data for all reservoirs and all time periods to define the threshold for determining extreme runoff for any reservoir and any time period. Using the extreme runoff threshold for any reservoir and any time period as the criterion, multiple sets of long-sequence simulated runoff of multiple reservoirs under normal scenarios are identified and labeled, and extreme scenario multi-reservoir runoff sequence data considering spatiotemporal two-dimensional correlation are obtained by filtering. Based on the runoff sequence data of multiple reservoirs under extreme scenarios, time-level cluster analysis was used to obtain single-period extreme scenarios and multi-period extreme scenarios, and spatial-level cluster analysis was used to obtain single-reservoir extreme scenarios and multi-reservoir extreme scenarios. The single-period extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff in any time period reaches the extreme threshold, it is defined as a single-period extreme abundance or a single-period extreme drought. The multi-period extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff in more than one time period reaches the extreme threshold, it is defined as multi-period extreme abundance, multi-period extreme scarcity, or multi-period abundance-scarcity transition. The extreme scenario for a single reservoir is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff of any reservoir reaches the extreme threshold, it is defined as an extreme abundance of a single reservoir, an extreme drought of a single reservoir, or a shift between abundance and drought in a single reservoir. The multi-reservoir extreme scenario is defined as follows: for a set of runoff sequences that include all reservoirs and all time periods, if the runoff of more than one reservoir reaches the extreme threshold, it is defined as multi-reservoir extreme abundance, multi-reservoir extreme drought, or multi-reservoir abundance-dry transition.
4. The method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff as described in claim 1, characterized in that, The optimization algorithm is a genetic algorithm, and its detailed implementation method includes: A1. Initialize the parameters of the genetic algorithm, using multiple sets of runoff sequence data from random simulations in step S2, including all reservoirs and all time periods, as model inputs; A2. Using the total discharge flow of all reservoirs, the average discharge flow over all time periods, or the reservoir capacity at the end of a time period as decision variables, randomly generate N different decision variables as the initial population. ; A3. Calculate the fitness of each individual in the population based on the objective function of the reservoir group optimization scheduling model; A4. Use the tournament selection algorithm to select individuals from the population to generate a population of size N / 2. ; A5. According to the preset crossover probability, select from the population Individuals are selected for crossover to generate offspring, and these offspring replace the original population. The parent individual in the middle; A6. Adjust the population according to the preset mutation probability. Individuals undergo mutation operations, and the mutated individuals replace the original individuals in the next generation of the population. Chromosomes that have not undergone mutation directly enter the next generation of the population. ; A7. Calculate the population based on the objective function of the reservoir group optimal scheduling model. The fitness value of each individual is calculated, and the current generation number Gen is incremented by 1; A8. Determine whether the number of generations Gen exceeds the set maximum number of generations or whether all fitness values meet the set threshold. If yes, proceed to step A9; otherwise, return to step A3. A9. Select the individual with the best fitness in the population as the optimal decision variable, and output the reservoir capacity at the end of the time period, the average discharge flow, the average power generation flow, the average water abandonment flow, and the average power output corresponding to the optimal decision variable for all reservoirs and all time periods, as the optimized scheduling result.
5. The method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff as described in claim 1, characterized in that, Step S4 further includes: A1. Constructing a multi-scheme sample set for data mining: Using multiple sets of runoff series data from multiple random simulations and corresponding multiple sets of reservoir group optimization scheduling calculation results, a multi-scheme sample set for data mining is constructed. Each scheme sample includes all reservoirs, initial reservoir capacity for all time periods, final reservoir capacity for all time periods, average inflow for each time period, average outflow for each time period, average power generation for each time period, average water discharge for each time period, and average power output for each time period. A2. Use data mining methods to identify key decision factors: B1. Calculate the grey relational degree of each decision factor. The average outflow or end-of-period reservoir capacity of all reservoirs is selected as the decision variable. The position of the period in the year, as well as the inflow, initial reservoir capacity, and average outflow of several previous and current periods of all reservoirs, are selected as candidate decision factors. The grey relational degree between the candidate decision factors and the decision variables is calculated one by one. The specific expression is as follows: in, Let be the grey relational degree of candidate decision factor j, with a value range of [0,1]. The grey relational coefficient of candidate decision factor j in record n; N is the total number of records; The value of the decision variable at record n; Let j be the value of candidate decision factor j at record n; The resolution coefficient; B2. Set the gray relational degree threshold Choose to satisfy The factors are used as key decision factors; A3. Construct a multi-scheme sample set for the machine learning model based on decision variables and key decision factors, and divide the multi-scheme sample set into a training set and a test set according to a preset ratio. Use the training set and the test set to train and validate the machine learning model to obtain the reservoir group scheduling rule model. The verification methods include using root mean square error, mean absolute error, and coefficient of determination as evaluation indicators to assess the effectiveness of the reservoir group scheduling rules of the machine learning method, and using reservoir capacity constraint over-limit rate and outflow over-limit rate as evaluation indicators to assess the feasibility of the reservoir group scheduling rules of the machine learning method.
6. The method for extracting reservoir group scheduling rules based on two-dimensional spatiotemporal stochastic simulation of runoff as described in claim 5, characterized in that, The machine learning model is the LightGBM algorithm, and a Bayesian optimization strategy is used to search for the optimal parameter combination of the LightGBM algorithm.