Machine learning based reservoir inflow runoff forecasting method, system and medium
By introducing the Multi-Objective Optimization Extreme Learning Machine (MOELM) model, combined with a fast elite genetic algorithm and sensitivity analysis, the problem of large prediction errors in existing models for extreme events is solved, achieving higher accuracy and adaptability in runoff prediction, and optimizing reservoir scheduling and power generation efficiency.
Patent Information
- Application Number
- CN202510791523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing data-driven models suffer from over-homogenization in runoff forecasting, especially when simulating extreme events, resulting in large forecast errors that affect reservoir management and increase risks. Furthermore, these models have poor adaptability and are difficult to apply effectively across different cross sections and scenarios.
A hybrid machine learning model based on Multi-Objective Optimization Extreme Learning Machine (MOELM) and combined with the fast elite multi-objective genetic algorithm NSGA-II was adopted to optimize the runoff simulation effect. Sensitivity analysis was performed using partial mutual information method (PMI) and empirical guidelines to improve the model's portability and accuracy.
The model improves the accuracy and reliability of runoff forecasting, reduces flood bias, enhances the safety of reservoir operation and the benefits of hydropower generation, and has strong adaptability and robustness among different reservoirs.
Smart Images

Figure CN120745898B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of hydrological prediction, specifically relating to a method, system, and medium for predicting reservoir inflow based on machine learning. Background Technology
[0002] Runoff forecasting plays a crucial role in water resource management and reservoir operation, especially during extreme weather events such as floods. Accurate forecasts are essential for optimizing reservoir operation, reducing risks, and improving hydropower efficiency. Existing data-driven models, particularly machine learning-based methods, have made some progress in runoff forecasting. However, a major problem with these models is over-homogenization, particularly when simulating extreme events. This can lead to excessive forecast errors, impacting reservoir operation decisions and subsequent management, and increasing potential risks.
[0003] To address this issue, recent years have seen an increasing number of studies attempting to improve forecast accuracy and reliability through hybrid models and optimization methods. However, traditional models still have limitations when dealing with complex flood events, failing to effectively reduce flood bias and exhibiting poor adaptability when transferred to different cross-sections and application scenarios. To overcome these challenges, this invention proposes a hybrid machine learning model based on Multi-Objective Optimization Extreme Learning Machine (MOELM). This model not only improves overall forecast accuracy but also focuses on the precise prediction of flood events, thereby optimizing reservoir operation and enhancing reservoir safety and hydropower efficiency. Summary of the Invention
[0004] The main objective of this invention is to overcome the shortcomings and deficiencies of the prior art and provide a method, system, and medium for predicting reservoir inflow based on machine learning, which can accurately predict the reservoir's next-day runoff and provide a more accurate method for runoff forecasting and analysis, thereby supporting reservoir scheduling.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In a first aspect, the present invention provides a method for predicting reservoir inflow based on machine learning, comprising the following steps:
[0007] S1. Collect reservoir inflow data and expand the runoff data to obtain a runoff sample set;
[0008] S2. Construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II). The objectives are to optimize the runoff simulation effect by minimizing the overall runoff fit and the root mean square error of extreme flood events. Specifically, this includes:
[0009] S21. Using the NSGA-II multi-objective optimization algorithm as a framework, randomly generate the connection weights IW and bias values B between the input layer and hidden layer of the ELM.
[0010] S22. Using the initially generated IW and B as parameters, construct an ELM model and use the historical inflow runoff series as input to obtain the simulated runoff;
[0011] S23. Calculate the relative error of the entire runoff series and the simulation error of the extreme flood runoff series based on simulated runoff and measured runoff;
[0012] S24. Based on the NSGA-II algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0013] S25. Output the IW and B obtained from the last iteration optimization of NSGA-II and the simulated runoff obtained from the ELM model to obtain the trained MOELM model.
[0014] S3. Sensitivity analysis of the MOELM model was conducted from the perspectives of model input information and number of hidden nodes using Partial Mutual Information (PMI) and empirical guidelines, respectively.
[0015] S4. Transfer the trained MOELM model to other reservoirs to verify the model's portability;
[0016] S5. The MOELM model, based on sensitivity analysis and portability test, is used to predict reservoir inflow.
[0017] As a preferred technical solution, in step S1, the inflow runoff sample set is expanded based on Markov model, specifically as follows:
[0018] The mean x and standard deviation S of the actual runoff samples were statistically analyzed. As upper and lower bounds, the runoff samples are divided into five groups from smallest to largest: dry year, relatively dry year, normal year, relatively abundant year, and abundant year runoff.
[0019] Establish state transition probability matrices for different categories of runoff samples and analyze the state transition probabilities.
[0020] Markov chains of discrete sequences can be represented by x. 2 To test using statistics, suppose the sequence under study contains m possible states, and use (f i,j Let i be the transition frequency probability matrix, j∈E. Divide the sum of each column of the transition frequency matrix by the sum of the sum of each row and column to obtain the "marginal probability" P. .j When m is large, x is analyzed using "marginal probability". 2 Statistic, at this time x 2 The statistic follows a sequence with (m-1) degrees of freedom.2 x 2 Given the distribution and confidence level α, x can be obtained by looking up the table. 2 >x α 2 ((m-1) 2 The value of ) when x 2 >x α 2 ((m-1) 2 ) represents x 2 The statistic is greater than the degrees of freedom (m-1). 2 Corresponding χ 2 The value refers to the value in χ². 2 In the distribution, given the confidence level α and degrees of freedom (m-1), 2 When the corresponding critical value is reached, the current runoff sample is considered to satisfy the "Madaspor property";
[0021] Calculate the autocorrelation coefficients of each order and solve for the weights of Markov chains with various step sizes;
[0022] To predict the required runoff samples, the weighted sum of the prediction probabilities for the same state is used as the index value, representing the prediction probability P of the i-th state. i ;
[0023] After determining the index value, add the index value to the original sequence, that is, complete the runoff sample supplementation for the i-th state. Repeat the above steps to supplement the remaining required samples, and finally obtain the runoff sample set.
[0024] Runoff samples in the runoff sample set are divided into training set, validation set and test set according to a set ratio, and used as input for training the MOELM model.
[0025] As a preferred technical solution, step S21 also includes the step of initializing the population using the NSGA-II multi-objective optimization algorithm, specifically as follows:
[0026] Population initialization: A certain number of individuals are randomly generated as the initial population, and each individual represents a potential solution, expressed as follows:
[0027]
[0028] Where x represents a single individual in NSGA-II. x The upper and lower bounds of the value range for this individual are defined by the connection between the weight IW and the bias value B, expressed as follows:
[0029]
[0030] Where ω n,1 Let b be the weight of the nth input layer and the weight of the 1st hidden layer in the input weights. n,1represents the bias values of the nth input layer and the 1st hidden layer.
[0031] As a preferred technical solution, step S22 specifically includes:
[0032] Based on the randomly generated initial weights IW and bias B, the input layer data is mapped from its original space to the feature space of the hidden layer through the activation function. Then, the output weights OW are determined according to the formula, and finally, the output layer ELM(X;L) is calculated, as follows:
[0033] ELM(X; L) = OW × g(IW × X + B)
[0034] ELM(X; L) represents simulated runoff; Q t X = [Q t-1 Q t-2 ,...,Q t-k [] is the input variable, representing historical inflow; IW, B, OW, and g(·) represent the input weights, bias, output weights, and activation function, respectively; L represents the number of hidden neurons; according to ELM theory, the simulated output matches the expected output, and the mathematical expression is:
[0035] ELM(X;L)=Q t
[0036] From a matrix perspective, the above equation can be rewritten as:
[0037] Q = Oβ
[0038] in:
[0039]
[0040] Output weight OW value It is uniquely determined using a formula, as follows:
[0041]
[0042] in The Moore-Penrose generalized inverse matrix representing β, once provided with new previous k observations, can be used with the training parameters IW, B, and To predict future runoff.
[0043] As a preferred technical solution, in step S23, the simulation errors of the entire runoff series and the extreme flood runoff series are expressed as follows:
[0044]
[0045] in, and They are respectively the i-thth The observed and simulated runoff from this observation set, where n is the total number of observations in the validation set; Q threshold The threshold for determining extreme flood events is when runoff data is greater than or equal to Q. threshold At that time, this flood was considered an extreme flood event. RMSE1 and RMSE2 represent the root mean square error (RMSE) of the total runoff and the extreme flood event, respectively, which are the simulation errors of the entire runoff series and the extreme flood runoff series.
[0046] As a preferred technical solution, the multi-objective optimization in step S24 specifically involves:
[0047] Non-dominated ordination: Based on the ELM model, the simulation of runoff aims to minimize the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. It uses non-dominated relationships to rapidly orient the population into different strata, as expressed below:
[0048] obj1 = minRMSE1
[0049] obj2=minRMSE2
[0050] Elite strategy: Retain the best solution from the previous generation as part of the next generation to expand the sampling space and prevent the loss of the best individuals;
[0051] Selection operation: Select individuals based on their fitness and non-dominance level, retaining superior individuals;
[0052] Mutation operation: After generating the initial population, a mutation operation is performed, which involves selecting a portion of the IW and B values for each target and mutating it. For each target x... i,G Let i = 1, 2, ..., NP, and let v be the mutation vector. i,G+1 The generation method is as follows:
[0053]
[0054] Among them, the randomly selected individual numbers r1, r2, r3 are required to be different from each other and also different from the target vector number i. Therefore, the population size must satisfy NP≥4. The mutation operator F∈[0,2] is a real constant factor that controls the scaling of the deviation variable.
[0055] From the perspective of runoff simulation, it can be expressed as:
[0056]
[0057] IW and B are the input weights and biases, respectively;
[0058] Crossover operation: Crossover operation is performed between the mutated population and the original population to generate a new generation of population (ui).j,G+1 The details are as follows:
[0059]
[0060] Where rand(0,1) represents a random value between [0,1], and CR represents the crossover operator, which takes values in the range of [0,1].
[0061] From the perspective of runoff simulation, it can be expressed as:
[0062]
[0063] Crowding comparison: Compare the crowding levels of individuals in adjacent levels and determine the survival ability of individuals based on the crowding level;
[0064] Iterative optimization: Repeat the above steps until the termination condition is met.
[0065] As a preferred technical solution, in step S3, sensitivity analysis of the MOELM model is performed from the perspectives of model input information and the number of hidden nodes using Partial Mutual Information (PMI) and empirical guidelines, respectively. Specifically:
[0066] S31. For sensitivity analysis, the partial mutual information (PMI) method is used, expressed as follows:
[0067] PMI(X,Y|Z)=D(p(x,y,z)||p * (x|z)p * (y|z)p(z))
[0068] Where p(x,y,z) is the joint probability distribution of random variables X, Y, Z, and D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) represents the path from p(x,y,z) to p * (x|z)p * The distance of (y|z)p(z);
[0069] In runoff forecasting simulation, for sensitivity analysis of the input layer structure, the partial mutual information method is first used to preliminarily screen out the set of forecast factors with the strongest correlation to runoff from the forecast factor set; then, based on the forecast factor combination obtained from the initial screening, the correlation coefficient method is used to calculate the correlation coefficient between the forecast factors and runoff at different lag times to identify the most important set of input forecast factors; the sensitivity is verified by analyzing the correlation coefficient and input importance of the selected factor set in the random forest (RF) model.
[0070] S32. For sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the hidden layer empirical guide, the input selected by PMI is used to test the number of hidden layer nodes in the MOELM model. The number of hidden nodes is changed in two intervals, and the runoff simulation effect under different hidden node numbers is analyzed. The empirical formula is expressed as follows:
[0071]
[0072] In the formula, γ is the number of neurons in the hidden layer, s is the number of runoff samples, n is the number of input layers, m is the number of output layer classifications, and k is a constant with a value in the range of [2, 10].
[0073] As a preferred technical solution, step S4 specifically includes:
[0074] S41. The objects of migration application include reservoirs with similar hydrogeological conditions in the same area as the research reservoir but different from them, as well as control sections of hydrological stations;
[0075] S42. The effectiveness of model transplantation was evaluated using four indicators: root mean square error, Nash efficiency coefficient, true positive rate, and false positive rate, as detailed below:
[0076] The root mean square error (RMSE) represents the average deviation between simulation and observation; a lower value indicates better model performance.
[0077] The Nash efficiency coefficient ranges between 1 and 0. According to the regulations of the Pearl River Resources Commission of the Ministry of Water Resources, the national hydrological forecasting standard requires NSE to be greater than 0.7, which is considered "acceptable".
[0078] True positive rate
[0079] TP represents the number of true positives, indicating that both predicted runoff and observed values exceed the threshold. This is a key number used to distinguish "flood" instances. FN represents the number of false negatives, indicating that the simulated operational event is not actually a "flood" event. A higher TPR can reduce operational risks.
[0080] False positive rate
[0081] FP stands for the number of false positives, or "false alarms," indicating that a flood was predicted but did not actually occur. TN stands for the number of true negatives, indicating that neither the forecast nor the observations predicted a flood. Lower FPR will increase the benefits of a reservoir because higher head contributes to hydropower generation.
[0082] Secondly, the present invention provides a reservoir inflow forecasting system based on machine learning, which is applied to the aforementioned machine learning-based reservoir inflow forecasting method, including a data acquisition module, a MOELM model construction module, a sensitivity analysis module, a model transfer module, and a runoff forecasting module.
[0083] The data acquisition module is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set.
[0084] The MOELM model building module is used to construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework, combined with the fast elite multi-objective genetic algorithm (NSGA-II), to optimize runoff simulation performance with the objectives of minimizing the overall runoff fit and the root mean square error of extreme flood events. Specifically, it includes:
[0085] Using the NSGA-II multi-objective optimization algorithm as a framework, the connection weights IW and bias values B between the input layer and hidden layer of the ELM are randomly generated;
[0086] Using the initially generated IW and B as parameters, an ELM model is constructed, and the simulated runoff is obtained by taking the historical inflow series as input.
[0087] The relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated based on simulated runoff and measured runoff.
[0088] Based on the NSGA-II algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0089] The output of IW and B obtained from the final iteration optimization of NSGA-II, and the simulated runoff obtained from the ELM model, will be used to obtain the trained MOELM model.
[0090] The sensitivity analysis module is used to perform sensitivity analysis on the MOELM model from the perspectives of model input information and number of hidden nodes, respectively, using Partial Mutual Information (PMI) and empirical guidelines.
[0091] The model porting module is used to transfer the trained MOELM model to other reservoirs to verify the model's portability.
[0092] The runoff forecasting module is used to forecast reservoir inflow based on the MOELM model, which is based on sensitivity analysis and portability testing.
[0093] Thirdly, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the aforementioned machine learning-based reservoir inflow forecasting method.
[0094] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0095] This invention introduces the MOELM hybrid machine learning method into the field of hydrological forecasting. It uses the Extreme Learning Machine (ELM) as the framework and combines it with the fast elite multi-objective genetic algorithm NSGA-II. The goal is to optimize the runoff simulation effect with the aim of minimizing the root mean square error of the overall runoff fit and extreme flood events. The model is analyzed from the perspective of the model input information and the number of hidden nodes through the partial mutual information (PMI) method.
[0096] Compared with traditional hydrological models, this method reduces operational risks and improves power generation efficiency by lowering the maximum outflow and water level in typical flood events. It also has strong hydrological migration capabilities, making it more scalable and robust in flood forecasting. This can effectively improve the accuracy of flood forecasting and reduce simulation errors for extreme floods. Attached Figure Description
[0097] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0098] Figure 1 This is a flowchart of a reservoir inflow forecasting method based on machine learning proposed in this invention;
[0099] Figure 2 This is a comparison chart of the overall performance of MOELM and the optimized learning machine OELM in flood event forecasting on a test set with the same time lag.
[0100] Figure 3 The MOELM performance results are illustrated using four typical flood events.
[0101] Figure 4 This is a comparison chart of the next-day runoff forecast results of this invention and the Xin'anjiang hydrological forecast model;
[0102] Figure 5 This is a comparison diagram of the results of applying the trained MOELM to other reservoirs and main cross sections;
[0103] Figure 6 This is a schematic diagram of the structure of the reservoir inflow forecasting system based on machine learning, as described in an embodiment of the present invention.
[0104] Figure 7 This is a schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present invention. Detailed Implementation
[0105] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0106] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0107] Please see Figure 1 This embodiment provides a reservoir inflow forecasting method based on machine learning, which includes the following steps:
[0108] S1. Initialize data and parameters.
[0109] Specifically, reservoir runoff data and related parameters were collected, including records of 632 flood events (flow exceeding 4000 m³) from January 1, 2008 to October 31, 2023, based on the operational time series. 3 Considering the lack of meteorological and hydrogeological data for the analyzed reservoir, the Markov method was used to enrich the inflow runoff sample series, and the runoff samples were divided based on the actual situation.
[0110] Furthermore, step S1 specifically includes:
[0111] S11. Collect the inflow data of the reservoir every six hours;
[0112] S12. To enrich the data series, the sample set is enriched based on Markov chains; the Markov method is used to expand the representation of inflow runoff samples as follows:
[0113] In Markov methods, the next state of a sequence depends only on the current state and not on past states. The formula is:
[0114] P(Y n+1 |Y1,Y2,...,Y n )=P(Y n+1 |Y n )
[0115] In the formula, Y1, Y2, ... are the state sequences in the Markov method. The equation shows that the sequence state at time n+1 is only related to time n.
[0116] The process of expanding runoff samples can be divided into the following steps:
[0117] First, the mean x and standard deviation S of the actual runoff samples are calculated. As upper and lower bounds, the runoff samples are divided into five groups from smallest to largest: dry year, relatively dry year, normal year, relatively abundant year, and abundant year runoff.
[0118] Secondly, state transition probability matrices for different categories of runoff samples are established, and the state transition probabilities are analyzed, expressed as follows:
[0119]
[0120] In the formula This represents the state transition probability of the k-th runoff sample from state i to state j; This represents the corresponding state transition probability matrix.
[0121] Furthermore, the Markov chain of discrete sequences can be represented by x. 2 To test using statistics, suppose the sequence under study contains m possible states, and use (f i,j Let E be the transition frequency probability matrix. Divide the sum of each column of the transition frequency matrix by the sum of the sum of each row and column to obtain the "marginal probability" P. .j When m is large, x is analyzed using "marginal probability". 2 Statistic, at this time x 2 The statistic follows a sequence with (m-1) degrees of freedom. 2 x 2 Given the distribution and confidence level α, x can be obtained by looking up the table. 2 >x α 2 ((m-1) 2 The value of ) when x 2 >x α 2 ((m-1) 2 (i.e., x) 2 The statistic is greater than the degrees of freedom (m-1). 2 Corresponding χ 2 "Value" refers to the value in χ². 2 In the distribution, with a given significance level (confidence level) α and degrees of freedom (m-1). 2 When the corresponding critical value is reached, the current runoff sample is considered to satisfy the "Madasaurus property", as expressed below:
[0122]
[0123] Then, calculate the autocorrelation coefficients r for each order. k Solve for the weights ω of Markov chains with various step sizes. k The expression is as follows:
[0124]
[0125] Finally, the required runoff samples are predicted. The weighted sum of the prediction probabilities for the same state is used as the index value, representing the prediction probability P for the i-th state. i The predicted probability is expressed as follows:
[0126]
[0127] After determining the index value, add it to the original sequence to supplement the runoff sample for state i. Repeat the above steps to supplement the remaining required samples.
[0128] S13. Following best practice, divide the runoff samples into model training, validation, and test sets at ratios of 70%, 15%, and 15%, and input them into the MOELM model for training.
[0129] S2. Construct and train a multi-objective optimized extreme learning machine (MOELM) model for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm NSGA-II (Nondominated Sorting Genetic Algorithm-II). The goal is to optimize the runoff simulation effect by minimizing the overall runoff fit and the root mean square error of extreme flood events. Specifically, this includes:
[0130] S21. Using the NSGA-II multi-objective optimization algorithm as a framework, randomly generate the connection weights IW and bias values B between the input layer and hidden layer of the ELM.
[0131] Furthermore, the specific steps for initializing the population in the NSGA-II multi-objective optimization algorithm are as follows:
[0132] Population initialization: A certain number of individuals are randomly generated as the initial population, and each individual represents a potential solution, expressed as follows:
[0133]
[0134] Where x represents a single individual in NSGA-II. x represents the upper and lower bounds of the value range for this individual. In this invention, it refers to the connection between the weight IW and the bias value B, expressed as follows:
[0135]
[0136] Where ω n,1 Let b be the weight of the nth input layer and the weight of the 1st hidden layer in the input weights. n,1 These are the bias values for the nth input layer and the 1st hidden layer.
[0137] S22. Using the initially generated IW and B as parameters, construct an ELM model and use the historical inflow runoff series as input to obtain the simulated runoff;
[0138] Furthermore, step S22 specifically includes:
[0139] Based on the randomly generated initial weights IW and bias B, the input layer data is mapped from its original space to the feature space of the hidden layer through the activation function. Then, the output weights OW are determined according to the formula, and finally, the output layer ELM(X;L) is calculated, as follows:
[0140] ELM(X; L) = OW × g(IW × X + B)
[0141] ELM(X;L) represents simulated runoff, Q t X = [Q t-1 Q t-2 ,...,Q t-k [] represents the input variable, i.e., historical inflow into the reservoir; IW, B, OW, and g(·) represent the input weights, bias, output weights, and activation function, respectively; L represents the number of hidden neurons. According to ELM theory, this model ideally approximates all training samples with non-zero errors, i.e., the simulated output matches the expected output. Its mathematical expression is:
[0142] ELM(X;L)=Q t
[0143] From a matrix perspective, the above equation can be written as:
[0144] Q = Oβ
[0145] in:
[0146]
[0147] The value of the output weight OW is uniquely determined using a formula (i.e. The details are as follows:
[0148]
[0149] in The Moore-Penrose generalized inverse matrix representing β, once provided with new previous k observations, can be used with the training parameters IW, B, and To predict future runoff.
[0150] S23. Calculate the relative error of the entire runoff series and the simulation error of the extreme flood runoff series based on simulated runoff and measured runoff;
[0151] Furthermore, the simulation errors for the entire runoff series and the extreme flood runoff series are expressed as follows:
[0152]
[0153] in and They are respectively the i-th th The observed and simulated runoff from this set of observations, where n is the total number of observations in the validation set. Q threshold The threshold for determining extreme flood events is set at which runoff data is considered an extreme flood event. The Q value for the Longtan Reservoir is... threshold 4000m 3 / s. RMSE1 and RMSE2 represent the root mean square error (RMSE) of total runoff and extreme flood events, respectively, which are the simulation errors of the entire runoff series and the extreme flood runoff series.
[0154] S24. Based on the NSGA-II algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0155] Furthermore, the specific methods for multi-objective optimization are as follows:
[0156] Non-dominated ordination: Based on the ELM model, the simulation of runoff aims to minimize the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. It uses non-dominated relationships to rapidly orient the population into different strata, as expressed below:
[0157] obj1 = minRMSE1
[0158] obj2=minRMSE2
[0159] Elite strategy: Retain the optimal solution from the previous generation as part of the next generation to expand the sampling space and prevent the loss of the best individuals.
[0160] Selection operation: Select individuals based on their fitness and non-dominant level, retaining the best individuals.
[0161] Mutation operation: After generating the initial population, a mutation operation is performed, which involves selecting a portion of the IW and B values for each target and mutating it. For each target x... i,G (i = 1, 2, ..., NP), the mutation vector is generated as follows:
[0162]
[0163] Among them, the randomly selected individual numbers r1, r2, r3 are required to be different from each other and also different from the target vector number i. Therefore, the population size must satisfy NP≥4. The mutation operator F∈[0,2] is a real constant factor that controls the scaling of the deviation variable.
[0164] From the perspective of runoff simulation, it can also be expressed as:
[0165]
[0166] Crossover operation: The mutated population is crossed with the original population to generate a new generation of population, as follows:
[0167]
[0168] Here, rand(0,1) represents a random value between [0,1], and CR represents the crossover operator, whose value range is [0,1].
[0169] From the perspective of runoff simulation, it can also be expressed as:
[0170]
[0171] Crowding comparison: Compare the crowding levels of individuals in adjacent levels and determine the survival ability of individuals based on the crowding level.
[0172] Iterative optimization: Repeat the above steps until the termination condition is met (such as reaching the maximum number of iterations or converging to a satisfactory Pareto front solution).
[0173] S25. Output the IW and B obtained from the last iteration optimization of NSGA-II and the simulated runoff obtained from the ELM model to obtain the trained MOELM model.
[0174] S3. Sensitivity analysis of the MOELM model was conducted from the perspectives of model input information and number of hidden nodes using Partial Mutual Information (PMI) and empirical guidelines, respectively.
[0175] Furthermore, step S3 specifically includes:
[0176] S31. For sensitivity analysis, the partial mutual information method (PMI) was selected, expressed as follows:
[0177] PMI(X,Y|Z)=D(p(x,y,z)||p * (x|z)p * (y|z)p(z))
[0178] Where p(x,y,z) is the joint probability distribution of random variables X, Y, Z, and D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) represents the path from p(x,y,z) to p * (x|z)p * The distance of (y|z)p(z).
[0179] In runoff forecasting simulation, for sensitivity analysis of the input layer structure, the partial mutual information method is first used to preliminarily screen out the set of forecast factors with the strongest correlation to runoff from the forecast factor set; then, based on the forecast factor combination obtained from the initial screening, the correlation coefficient method is used to calculate the correlation coefficient between the forecast factors and runoff at different lag times to identify the most important set of input forecast factors; the sensitivity is verified by analyzing the correlation coefficient and input importance of the selected factor set in the random forest (RF) model.
[0180] S32. For the sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the hidden layer empirical guide, the input selected by PMI is used to test the number of hidden layer nodes in the MOELM model. The number of hidden nodes is changed in two intervals, and the runoff simulation effect under different hidden node numbers is analyzed. The empirical formula is expressed as follows:
[0181]
[0182] In the formula, γ is the number of neurons in the hidden layer, s is the number of runoff samples, n is the number of input layers, m is the number of output layer classifications, and k is a constant with a value in the range of [2, 10].
[0183] Table 1 lists the maximum PMI value (from mutual information in the initial round) and the 95% confidence limit for each round. The correlation coefficients and input importance rankings of the selected inputs with the Random Forest (RF) model are good, indicating that these inputs were indeed significant when selected.
[0184] However, the selected candidate inputs overlap considerably with those in the lag 1–4 scenarios, which could lead to slight differences in evaluation metrics. To address this, a lag 4–7 scenario was proposed to demonstrate the performance of four equivalent inputs. The results show that the inputs selected by PMI consistently achieve the best performance across all four input scenarios. Specifically, compared to the lag 4–7 scenario, the improvements in overall RMSE1, RMSE2, NSE, TPR, and FPR are 12.7%, 35.5%, 8.9%, 29.2%, and 32.2%, respectively, and compared to the lag 1–4 scenario, they are 1.1%, 1.4%, 5.6%, 2.8%, and 13.2%, respectively. These results validate the intuition that input variables should be selected. Although the differences between PMI and the lag 1–4 scenarios are not significant, the selected input variables generally align with the most relevant input variables in the sequence.
[0185] Table 1. Stepwise predictor selection with partial mutual information (PMI)
[0186]
[0187] S4. Transfer the trained MOELM model to other reservoirs to verify the model's portability;
[0188] Furthermore, the portability of the model is specifically tested as follows:
[0189] The S8-1 migration application targets include reservoirs in the same area as the research reservoir but different from them, as well as control sections of hydrological stations.
[0190] The effectiveness of the S8-2 model transplantation application was evaluated using four indicators: root mean square error, Nash efficiency coefficient, true positive rate, and false positive rate, as detailed below:
[0191] Root Mean Square Error (RMSE)
[0192]
[0193] RMSE represents the average deviation between simulation and observation; a lower value indicates better performance. In the formula... Let i be the simulated runoff volume for the i-th time period. Let be the observed runoff volume for the i-th time period.
[0194] Nash-Sutcliffe efficiency coefficient (NSE)
[0195]
[0196] in This represents the average of the observed values. The NSE ranges between 1 (perfect fit) and 0. According to the regulations of the Pearl River Resources Commission of the Ministry of Water Resources, the national hydrological forecasting standard requires an NSE greater than 0.7 to be considered "acceptable".
[0197] True positive rate (TPR)
[0198]
[0199] TP represents the number of true positives, indicating that both predicted runoff and observed values exceed the threshold; this is a key figure for distinguishing instances of “flooding.” FN represents the number of false negatives, indicating that the simulated operational event is not actually a “flooding” event. A higher TPR can reduce operational risk, as unexpected floods can burden operational plans (Huang et al. 2022).
[0200] False positive rate (FPR)
[0201]
[0202] FP stands for the number of false positives, or "false alarms," indicating that a flood was predicted but did not actually occur. TN stands for the number of true negatives, indicating that neither the forecast nor observations predicted a flood. A lower FPR will increase the benefits of a reservoir because a higher head contributes to hydroelectric power generation.
[0203] S5. The MOELM model, based on sensitivity analysis and portability test, is used to predict reservoir inflow.
[0204] The results of the optimized Extreme Learning Machine (OELM) simulation forecast of the flood event in this reservoir were compared with those of the actual flood. The results are as follows: Figure 2 As shown, MOELM achieved a higher prediction accuracy than OELM. The average improvement was approximately 5.27% across scenarios with different input delay times. The initial improvement was not significant, but it steadily increased, highlighting the importance of input selection. While MOELM slightly increased the RMSE across all records, this was negligible at a marginal rate of around 0.9%, and the NSE value only slightly decreased from 0.8 to 0.79. This difference is acceptable for subsequent reservoir management under extreme events; the four types of measured flood events selected were as follows: two peak flows of approximately 10,000 m³ / h. 3 / s represents a 20-year recurrence period, with both peak flows estimated at approximately 6000 m³ / s. 3 / s, representing a 10-year return period, became a reliable flood forecast consistent with observations. The deviation of peak flow was less than 5%.
[0205] like Figure 3As shown, compared with the simulation results of Linear Regression (LR), Single-Layer Feedforward Network (SLFN), Extreme Gradient Boosting (XGboost), and Random Forest (RF), MOELM generally achieved the best performance in all tests. Table 2 lists the average values of all 20 runs with different lag times. The results show that the algorithms based on RMSE, NSE, and TPR differ significantly based on the group t-test results at p=0.05. However, FPR shows the opposite trend, meaning that the optimal algorithm did not achieve the minimum FPR. The false alarm rate in the MOELM model is quite low, approximately 0.2%.
[0206] Table 2 shows the average running time after different methods.
[0207]
[0208]
[0209]
[0210] By comparing the simulation results of the Xin'anjiang model, reviewing the observed historical runoff, and analyzing the results regarding maximum water level, maximum reservoir output, and hydropower operation risks, the results are as follows: Figure 4 As shown, the maximum flow rate increased by an average of 490m in the relevant scenario calculations. 3 / s, with an average water level increase of 1.3m, the risks associated with MOELM are negligible. For hydropower, using MOELM's predicted flow, hydropower plants can generate an additional 130 million kilowatt-hours.
[0211] To evaluate the transferability of the model, the invention was applied to the control sections of three large reservoirs (Tianyi, Baise, and Changzhou) and the Wuzhou station in the Pearl River Basin. The results are as follows: Figure 5 As shown, the model exhibits excellent transferability. The NSE values for Tianyi and Baise reservoirs are 0.75 and 0.65, respectively, which are 6.5% and 18.7% lower than those for Longtan reservoir. The meteorological and hydrological conditions of Tianyi reservoir are similar to those of Longtan reservoir, resulting in less variation. The conditions of Baise reservoir are significantly different, leading to a substantial decrease in model performance. The model performs exceptionally well in Wuzhou and Changzhou, with NSE values exceeding 0.95. This superior performance is attributed to variations in flow direction. The rates of variation (relative to the average level) in Wuzhou and Changzhou are approximately 8.5% and 4.7%, respectively, both significantly lower than the 56.3% of Longtan reservoir.
[0212] This invention can accurately predict the next day's runoff from a reservoir, providing a more precise method for runoff forecasting and analysis.
[0213] It should be noted that, for the sake of simplicity, the aforementioned method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously.
[0214] Based on the same idea as the machine learning-based reservoir inflow forecasting method in the above embodiments, this invention also provides a machine learning-based reservoir inflow forecasting system, which can be used to execute the above-described machine learning-based reservoir inflow forecasting method. For ease of explanation, the structural diagram of the machine learning-based reservoir inflow forecasting system embodiment only shows the parts related to the embodiments of this invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0215] Please see Figure 6 In another embodiment of this application, a reservoir inflow forecasting system 100 based on machine learning is provided. The system includes a data acquisition module 101, a MOELM model building module 102, a sensitivity analysis module 103, a model transfer module 104, and a runoff forecasting module 105.
[0216] The data acquisition module 101 is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set.
[0217] The MOELM model construction module 102 is used to construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework, combined with the fast elite multi-objective genetic algorithm NSGA-II, to optimize runoff simulation performance with the objectives of minimizing the overall runoff fit and the root mean square error of extreme flood events. Specifically, it includes:
[0218] Using the NSGA-II multi-objective optimization algorithm as a framework, the connection weights IW and bias values B between the input layer and hidden layer of the ELM are randomly generated;
[0219] Using the initially generated IW and B as parameters, an ELM model is constructed, and the simulated runoff is obtained by taking the historical inflow series as input.
[0220] The relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated based on simulated runoff and measured runoff.
[0221] Based on the NSGA-II algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0222] The output of IW and B obtained from the final iteration optimization of NSGA-II, and the simulated runoff obtained from the ELM model, will be used to obtain the trained MOELM model.
[0223] The sensitivity analysis module 103 is used to perform sensitivity analysis on the MOELM model from the perspectives of model input information and number of hidden nodes using Partial Mutual Information (PMI) and empirical guidelines, respectively.
[0224] The model porting module 104 is used to transfer the trained MOELM model to other reservoirs to verify the model's portability.
[0225] The runoff forecasting module 105 is used to forecast the inflow runoff into the reservoir based on the MOELM model, which is based on sensitivity analysis and portability test.
[0226] It should be noted that the reservoir inflow forecasting system based on machine learning of the present invention corresponds one-to-one with the reservoir inflow forecasting method based on machine learning of the present invention. The technical features and beneficial effects described in the embodiments of the reservoir inflow forecasting method based on machine learning are applicable to the embodiments of reservoir inflow forecasting based on machine learning. For details, please refer to the description in the embodiments of the method of the present invention, which will not be repeated here.
[0227] Furthermore, in the above embodiments of the reservoir inflow forecasting system based on machine learning, the logical division of each program module is only an example. In actual applications, the above functions can be assigned to different program modules as needed, for example, for the sake of corresponding hardware configuration requirements or software implementation convenience. That is, the internal structure of the reservoir inflow forecasting system based on machine learning can be divided into different program modules to complete all or part of the functions described above.
[0228] Please see Figure 7 In one embodiment, this embodiment provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the machine learning-based reservoir inflow forecasting method, specifically:
[0229] S1. Collect reservoir inflow data and expand the runoff data to obtain a runoff sample set;
[0230] S2. Construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II). The objectives are to optimize the runoff simulation effect by minimizing the overall runoff fit and the root mean square error of extreme flood events. Specifically, this includes:
[0231] S21. Using the NSGA-II multi-objective optimization algorithm as a framework, randomly generate the connection weights IW and bias values B between the input layer and hidden layer of the ELM.
[0232] S22. Using the initially generated IW and B as parameters, construct an ELM model and use the historical inflow runoff series as input to obtain the simulated runoff;
[0233] S23. Calculate the relative error of the entire runoff series and the simulation error of the extreme flood runoff series based on simulated runoff and measured runoff;
[0234] S24. Based on the NSGA-II algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0235] S25. Output the IW and B obtained from the last iteration optimization of NSGA-II and the simulated runoff obtained from the ELM model to obtain the trained MOELM model.
[0236] S3. Sensitivity analysis of the MOELM model was conducted from the perspectives of model input information and number of hidden nodes using Partial Mutual Information (PMI) and empirical guidelines, respectively.
[0237] S4. Transfer the trained MOELM model to other reservoirs to verify the model's portability;
[0238] S5. The MOELM model, based on sensitivity analysis and portability test, is used to predict reservoir inflow.
[0239] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0240] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0241] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for reservoir inflow runoff forecasting based on machine learning, characterized by, The method comprises the following steps: S1, collecting reservoir inflow runoff data, and expanding the runoff data to obtain a runoff sample set; S2, constructing a multi-objective optimization extreme learning machine model MOELM for reservoir inflow runoff prediction and training the model, the MOELM model is based on an extreme learning machine ELM as a framework, combined with a fast elitist multi-objective genetic algorithm NSGA-II, and the root mean square error of overall runoff fitting and extreme flood events is minimized to optimize the runoff simulation effect, specifically comprising: S21, taking the NSGA-II multi-objective optimization algorithm as a framework, randomly generating connection weights IW and bias weights B between the input layer and the hidden layer of the ELM; S22, taking the initially generated IW and B as parameters, constructing an ELM model, and taking a historical inflow runoff series as input to obtain simulated runoff; S23, calculating the relative error of the entire runoff series and the simulation error of the extreme flood runoff series based on the simulated runoff and the measured runoff; S24, based on the NSGA-II algorithm, minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series as the target, and performing multi-objective optimization; S25, taking the IW and B obtained by the last iteration optimization of the NSGA-II and the simulated runoff obtained according to the ELM model as input to obtain a trained MOELM model; S3, performing sensitivity analysis on the MOELM model from the angles of model input information and hidden node number through the partial mutual information method PMI and the empirical guideline; S4, migrating the trained MOELM model to other reservoirs to test the portability of the model; S5, predicting the reservoir inflow runoff based on the sensitivity analysis and the portability test of the MOELM model; In the step S1, the inflow runoff sample set is expanded based on the Markov method, specifically: Statistical mean of actual runoff samples and standard deviation ,by As upper and lower bounds, the runoff samples are divided into five groups from smallest to largest: dry year, relatively dry year, normal year, relatively abundant year, and abundant year runoff. A state transition probability matrix of different types of runoff samples is established, and the state transition probability is analyzed; Markov chains of discrete sequences can be represented by x. 2 To test using statistics, suppose the sequence under study contains m possible states, use... Let E be the transition frequency probability matrix, i, j∈E. Divide the sum of each column of the transition frequency probability matrix by the sum of the sum of each row and column to obtain the "marginal probability". Analysis using "marginal probability" Statistic, at this time The statistic follows a sequence with (m-1) degrees of freedom. 2 of Distribution, giving confidence level From the table, we can obtain The value when express The statistic is greater than the degrees of freedom (m-1). 2 Corresponding χ 2 The value refers to the value in χ². 2 In the distribution, given the confidence level α and degrees of freedom (m−1)... 2 When the corresponding critical value is reached, the current runoff sample is considered to satisfy the "Madaspor property"; The autocorrelation coefficients of each order are calculated, and the weights of Markov chains of various step lengths are solved; The prediction probability of the same state is weighted and summed as an index value at the prediction probability of the state ; After the index value is determined, the index value is added to the original sequence, that is, the first The runoff sample set is obtained by supplementing the state of the runoff sample and supplementing the remaining required samples. The runoff samples in the runoff sample set are divided into a training set, a validation set and a test set according to a set proportion, which are used as input for the MOELM model training; In the step S3, the sensitivity analysis of the MOELM model is performed from the angles of model input information and hidden node number through the partial mutual information method PMI and the empirical guideline, specifically: S31, for sensitivity analysis, the partial mutual information method PMI is selected, which is expressed as follows: wherein, the joint probability distribution, denotes the distance from the distance from In runoff prediction simulation, for input layer structure sensitivity analysis, the partial mutual information method is first used to preliminarily screen out the strongest runoff correlation factor set from the prediction factor set; then, according to the prediction factor combination obtained by preliminary screening, the correlation coefficient method is used to calculate the correlation coefficient of the prediction factor and the runoff under different lags, and the most important input prediction factor set is identified; the sensitivity is verified by analyzing the correlation coefficient and input importance of the selected factor set in the random forest RF model. S32, for the sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the empirical guide of the hidden layer, the input selected by the PMI is used to test the number of hidden layer nodes of the MOELM model, the number of hidden nodes is changed in two intervals, the runoff simulation effect under different number of hidden nodes is analyzed, and the empirical formula is expressed as follows: wherein, is the number of hidden layer neurons, is the number of runoff samples, is the number of input layers, is the number of output layer classifications, is a constant with a value range between [2, 10]. 2.The method of claim 1, wherein, The step S21 further includes the step of initializing the population of the NSGA-II multi-objective optimization algorithm, specifically: Initializing the population: a certain number of individuals are randomly generated as the initial population, and each individual represents a potential solution, which is expressed as follows: wherein is the upper bound of the value range of the individual, , is the lower bound of the value range of the individual, then the connection weight and the bias value are expressed as follows: in For the first The input layer, and the weights of the first hidden layer in the input weights. represents the bias values of the nth input layer and the 1st hidden layer. 3.The method of claim 1, wherein, The step S22 is specifically: According to the randomly generated initialization weights and bias , the input layer data is mapped from its original space to the feature space of the hidden layer by the activation function, and then the output weight is determined according to the formula , and finally the output of the output layer is calculated , as follows: denotes the analog runoff; , is an input variable denoting historical inflow runoff; , , and denote input weights, bias, output weights, and activation function, respectively; denotes the number of hidden neurons; according to the ELM theory, the analog output matches the expected output, and the mathematical expression is: From the perspective of matrix, the above formula is rewritten as: Wherein: output weight the value is uniquely determined using the formula, as follows: wherein represent the Moore-Penrose generalized inverse matrix, once the new previous k observations are provided, the trained parameters , and can be used to forecast future runoff. 4.The method of claim 1, wherein, In the step S23, the simulation error of the whole runoff series and the extreme flood runoff series is expressed as follows: where, and are the observed and simulated runoff of the first observation, is the total number of observations in the validation set; is the threshold value for extreme flood events, when the runoff data is greater than or equal to the flood is considered an extreme flood event, and are the root mean square errors of the total runoff and extreme flood events i.e. the simulation error of the entire runoff series and the extreme flood runoff series. 5.The method of claim 4, wherein, In the step S24, the multi-objective optimization is specifically: Non-dominated sorting: based on the simulated runoff of the ELM model, the relative error of the whole runoff series and the simulation error of the extreme flood runoff series are minimized as the target, the population is quickly non-dominated sorted according to the non-dominated relationship, and the population is divided into different levels, and the target is expressed as follows: Elite strategy: the optimal solution of the last generation is reserved as part of the next generation, so as to expand the sampling space and prevent the loss of the best individual; Selection operation: selection operation is performed according to the fitness and non-dominated level of the individual, and excellent individuals are reserved; Mutation operation: After the initial population is generated, mutation operation is performed, i.e. the values in the selected part of each objective are mutated, for each objective and The mutation is performed on the selected part of the value selection part, for each objective The mutation vector is generated as follows: where the randomly selected individual index are different, and the target vector index are also different, so the population size must satisfy , the mutation operator is a real constant factor, controlling the scaling of the perturbation variable; From the perspective of runoff simulation, it is expressed as: IW and B are input weight and bias respectively; Crossover operation: the mutated population is crossed with the initial population to generate a new generation of population , as follows: Wherein, rand(0,1) represents a random value between [0,1], CR represents a crossover operator, and the value range is [0,1]; From the perspective of runoff simulation, it is expressed as: Crowded comparison: compare the crowded degree of individuals in adjacent levels, and determine the survival ability of the individual according to the crowded degree; Iterative optimization: until the termination condition is met. 6.The method of claim 1, wherein, The step S4 is specifically: S41, the migration application object includes different reservoirs with similar hydrogeological conditions in the same region as the research reservoir, and the control section of the hydrological station; S42, the model transplantation application effect selects four indexes of root mean square error, Nash efficiency coefficient, true positive rate and false positive rate for evaluation, which are specifically as follows: The root mean square error represents the average deviation between simulation and observation, and the lower the value, the better the model performance; The Nash efficiency coefficient ranges between 1 and 0, according to the regulations of the Pearl River Resources Commission of the Ministry of Water Resources, the national hydrological forecasting standards require greater than 0.7, "acceptable"; True positive rate ; where is the number of true positives, indicating that both the forecasted runoff and the observed value exceeded the threshold, which is the key number for distinguishing "flood" instances; is the number of false negatives, indicating that the simulated run event was not actually a "flood" event; False positive rate ; wherein is the number of false positives, or "false alarms," indicating that a flood was forecast, but no flood actually occurred, is the number of true negatives, indicating that neither the forecast nor the observations forecast a flood.
7. A machine learning based reservoir inflow runoff forecasting system characterized in that, The machine learning-based reservoir inflow runoff prediction method according to any one of claims 1-6 comprises a data acquisition module, a MOELM model construction module, a sensitivity analysis module, a model transplantation module and a runoff prediction module; The data acquisition module is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set; The MOELM model construction module is used to construct and train a multi-objective optimization extreme learning machine model MOELM for reservoir inflow runoff prediction, the MOELM model is based on an extreme learning machine ELM as a framework, combined with a fast elite multi-objective genetic algorithm NSGA-II, and the root mean square error of the whole runoff fitting and the extreme flood event is minimized as the target to optimize the runoff simulation effect, which specifically includes: Taking NSGA-Ⅱ multi-objective optimization algorithm as a framework, connection weights IW and bias weights B between input layer and hidden layer of ELM are randomly generated; Taking the initial generated IW and B as parameters, an ELM model is constructed, and simulated runoff is obtained by taking historical reservoir inflow series as input; Based on the simulated runoff and the measured runoff, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated; Based on the NSGA-Ⅱ algorithm, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series are minimized as the target, and multi-objective optimization is performed; The output IW and B obtained by the last iteration optimization of NSGA-Ⅱ and the simulated runoff obtained according to the ELM model are used to obtain the trained MOELM model; The sensitivity analysis module is configured to perform sensitivity analysis on the MOELM model from the perspectives of model input information and hidden node number by using the partial mutual information method (PMI) and the empirical guideline, respectively. The model transplantation module is configured to migrate the trained MOELM model to other reservoirs to verify the portability of the model. The runoff prediction module is configured to predict the reservoir inflow runoff based on the MOELM model that has been subjected to sensitivity analysis and portability verification.
8. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, the machine learning-based reservoir inflow runoff prediction method of any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Flood control optimization scheduling method based on flood type self-identification strategy
CN116502773A
Port rainfall runoff simulation and waterlogging forecasting method and system based on SWMM
CN119294297A