Reservoir inflow runoff forecasting method and system based on machine learning and medium
By introducing the MOELM model and NSGA-Ⅱ algorithm to optimize runoff forecast, the problem of large errors in extreme event forecasts in existing models is solved, and runoff forecasts with higher accuracy and greater adaptability are achieved, thereby improving the efficiency of reservoir scheduling and power generation.
Patent Information
- Application Number
- CN202510791523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing data-driven models suffer from over-homogenization in runoff forecasting, especially when simulating extreme events, resulting in large forecast errors, which affect reservoir scheduling and increase risks. In addition, the models have poor adaptability and are difficult to effectively apply in different sections and scenarios.
A hybrid machine learning model based on multi-objective optimization extreme learning machine (MOELM) is adopted, combined with the fast elite multi-objective genetic algorithm NSGA-Ⅱ, to optimize the runoff simulation effect. The sensitivity of model input information and the number of hidden nodes is analyzed through partial mutual information method PMI and empirical guidelines to improve the forecast accuracy of extreme flood events.
The overall accuracy of runoff forecasts and the accuracy of extreme events have been improved, flood deviations have been reduced, the safety of reservoir scheduling and the benefits of hydropower generation have been enhanced, and the model has strong migration capabilities and robustness among different reservoirs.
Smart Images

Figure CN120745898A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hydrological forecasting, and specifically relates to a reservoir inflow runoff forecasting method, system and medium based on machine learning. Background Art
[0002] Runoff forecasting plays a crucial role in water resource management and reservoir operation. Accurate forecasts are particularly important for optimizing reservoir operation, reducing risks, and improving hydropower generation efficiency during extreme weather events, such as floods. Existing data-driven models have made some progress in runoff forecasting, particularly those based on machine learning. However, a major problem with existing models is over-homogenization, which can lead to large forecast errors, particularly when simulating extreme events. This can affect reservoir operation decisions and subsequent operations, increasing potential risks.
[0003] To address this problem, in recent years, more and more studies have attempted to improve the accuracy and reliability of forecasts through hybrid models and optimization methods. However, traditional models still have limitations when dealing with complex flood events, cannot effectively reduce flood deviations, and have poor adaptability when migrating to different sections and application scenarios. In order to overcome these challenges, the present invention proposes a hybrid machine learning model based on a multi-objective optimization extreme learning machine (MOELM). This model not only improves the overall forecast accuracy, but also pays more attention to the accurate forecast of flood events, thereby optimizing reservoir scheduling and improving the safety of reservoirs and the benefits of hydropower generation. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology, and provide a reservoir inflow runoff forecasting method, system and medium based on machine learning, accurately forecast the reservoir runoff for the next day, and provide a more accurate method for runoff forecast analysis, thereby providing support for reservoir scheduling.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for predicting reservoir inflow runoff based on machine learning, comprising the following steps:
[0007] S1. Collect reservoir runoff data and expand the runoff data to obtain a runoff sample set;
[0008] S2. Construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II). It optimizes the runoff simulation effect by minimizing the root mean square error of the overall runoff fit and extreme flood events. Specifically, it includes:
[0009] S21. Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated;
[0010] S22, using the initially generated IW and B as parameters, construct the ELM model, and use the historical inflow runoff series as input to obtain the simulated runoff;
[0011] S23, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series calculated based on the simulated runoff and the measured runoff;
[0012] S24. Based on the NSGA-Ⅱ algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series;
[0013] S25, outputting IW and B obtained by the last iterative optimization of NSGA-II and the simulated runoff obtained according to the ELM model to obtain a trained MOELM model;
[0014] S3, sensitivity analysis of the MOELM model was performed from the perspectives of model input information and number of hidden nodes using partial mutual information method (PMI) and empirical guidelines;
[0015] S4, migrate the trained MOELM model to other reservoirs to test the portability of the model;
[0016] S5. The MOELM model is used to predict reservoir inflow based on sensitivity analysis and portability test.
[0017] As a preferred technical solution, in step S1, the runoff sample set is expanded based on Markov, specifically:
[0018] Statistical actual runoff sample mean x and standard deviation S, As the upper and lower bounds, the runoff samples are divided into five groups from small to large: dry year, relatively dry year, normal year, relatively wet year, and wet year runoff;
[0019] Establish the state transition probability matrix of different types of runoff samples and analyze the state transition probability;
[0020] The Markov chain of discrete sequences can be expressed as x 2 Statistics to test, suppose the sequence under study contains m possible states, use (f i,j )i is recorded as the transfer frequency probability matrix, j∈E, and the sum of each column of the transfer frequency matrix is divided by the sum of each row and column to obtain the "marginal probability" P .j , when m is large, analyze x with "marginal probability" 2 Statistics, at this time x 2 The statistic is subject to degrees of freedom (m-1)2 x 2 Distribution, given the confidence α, we can get x by looking up the table 2 >x α 2 ((m-1) 2 ) value, when x 2 >x α 2 ((m-1) 2 ) represents x 2 The statistic is greater than the degrees of freedom (m-1) 2 The corresponding χ 2 The value refers to the 2 distribution, with a given confidence level α and degrees of freedom (m-1) 2 When the corresponding critical value is reached, the current runoff sample is considered to meet the “Markov property”;
[0021] Calculate the autocorrelation coefficients of each order and solve the weights of Markov chains with various step sizes;
[0022] Predict the required runoff samples, and take the weighted sum of the prediction probabilities of the same state as the prediction probability P of the index value in the i-th state i ;
[0023] After determining the index value, add the index value to the original sequence, that is, realize the runoff sample supplement of the i-th state, repeat the above steps to supplement the remaining required samples, and finally obtain the runoff sample set;
[0024] The runoff samples in the runoff sample set are divided into training set, validation set and test set according to the set ratio, which serve as the input for MOELM model training.
[0025] As a preferred technical solution, step S21 also includes the step of initializing the population using the NSGA-II multi-objective optimization algorithm, specifically:
[0026] Initialize the population: Randomly generate a certain number of individuals as the initial population. Each individual represents a potential solution, which can be expressed as follows:
[0027]
[0028] Where x is a single individual in NSGA-Ⅱ, x The upper and lower bounds of the individual value range refer to the connection weight IW and the bias value B, which are expressed as follows:
[0029]
[0030] where ω n,1 is the nth input layer, the weight of the first hidden layer in the input weight, b n,1is the bias value of the nth input layer and the first hidden layer.
[0031] As a preferred technical solution, step S22 is specifically as follows:
[0032] According to the randomly generated initialization weights IW and bias B, the input layer data is mapped from its original space to the feature space of the hidden layer through the activation function, and then the output weight OW is determined according to the formula, and finally the output layer output ELM(X; L) is obtained, as follows:
[0033] ELM(X;L)=OW×g(IW×X+B)
[0034] ELM(X;L) represents simulated runoff; Q t , X=[Q t-1 ,Q t-2 ,...,Q t-k ] is the input variable, representing the historical inflow runoff; IW, B, OW, and g(·) represent the input weight, bias, output weight, and activation function, respectively; L represents the number of hidden neurons; According to the ELM theory, the simulated output matches the expected output, and the mathematical expression is:
[0035] ELM(X;L)=Q t
[0036] From the perspective of the matrix, the above formula can be rewritten as:
[0037] Q=Oβ
[0038] in:
[0039]
[0040] Output weight OW value It is uniquely determined using the formula, as follows:
[0041]
[0042] in The Moore-Penrose generalized inverse matrix of β is used to train the parameters IW, B, and To predict future runoff.
[0043] As a preferred technical solution, in step S23, the simulation errors of the entire runoff series and the extreme flood runoff series are expressed as follows:
[0044]
[0045] in, and The ith The observed and simulated runoff is n observations, n is the total number of observations in the validation set; Q threshold is the threshold for determining extreme flood events. When the runoff data is greater than or equal to Q threshold When , this flood is considered to be an extreme flood event. RMSE1 and RMSE2 represent the root mean square error RMSE of the total runoff and the extreme flood event, respectively, that is, the simulation errors of the entire runoff series and the extreme flood runoff series.
[0046] As a preferred technical solution, in step S24, the multi-objective optimization is specifically performed as follows:
[0047] Non-dominated sorting: Based on the ELM model, the runoff simulation aims to minimize the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. The population is quickly sorted according to the non-dominated relationship and the population is divided into different levels. The goal is expressed as follows:
[0048] obj1=minRMSE1
[0049] obj2=minRMSE2
[0050] Elite strategy: retain the best solution of the previous generation as part of the next generation to expand the sampling space and prevent the loss of the best individuals;
[0051] Selection operation: Perform selection operations based on individual fitness and non-dominated hierarchy to retain excellent individuals;
[0052] Mutation operation: After the initial population is generated, the mutation operation is performed, that is, the IW and B value selection parts in each target are mutated. For each target x i,G , i=1,2,...,NP, mutation vector v i,G+1 The generation method is as follows:
[0053]
[0054] Among them, the randomly selected individual serial numbers r1, r2, r3 are required to be different from each other and also different from the target vector serial number i, so the population size must satisfy NP ≥ 4, and the mutation operator F∈[0,2] is a real constant factor that controls the scaling of the deviation variable;
[0055] From the perspective of runoff simulation, it can be expressed as:
[0056]
[0057] IW and B are input weight and bias respectively;
[0058] Crossover operation: Cross the mutated population with the initial population to generate a new generation of population uij,G+1 , as follows:
[0059]
[0060] Among them, rand(0,1) represents a random value between [0,1], CR represents the crossover operator, and its value range is [0,1];
[0061] From the perspective of runoff simulation, it can be expressed as:
[0062]
[0063] Crowding comparison: Compare the crowding levels of individuals in adjacent levels and determine the survival ability of individuals based on the crowding level;
[0064] Iterative optimization: Repeat the above steps until the termination condition is met.
[0065] As a preferred technical solution, in step S3, sensitivity analysis of the MOELM model is performed from the perspectives of model input information and the number of hidden nodes using the partial mutual information method PMI and empirical guidance, specifically:
[0066] S31. For sensitivity analysis, the partial mutual information method PMI is selected, which is expressed as follows:
[0067] PMI(X,Y|Z)=D(p(x,y,z)||p * (x|z)p * (y|z)p(z))
[0068] Where p(x,y,z) is the joint probability distribution of random variables X,Y,Z, and D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) means p(x,y,z) to p * (x|z)p * The distance of (y|z)p(z);
[0069] In runoff forecast simulation, for the sensitivity analysis of the input layer structure, the partial mutual information method is first used to preliminarily screen out the set of predictors with the strongest correlation with runoff from the set of predictors. Then, based on the predictor combinations obtained from the initial screening, the correlation coefficient method is used to calculate the correlation coefficients between the predictors and runoff at different lags to identify the most important input predictor sets. The sensitivity is verified by analyzing the correlation coefficients and input importance of the selected factor sets in the random forest RF model.
[0070] S32. For the sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the hidden layer experience guide, the input selected by PMI was used to test the number of hidden layer nodes of the MOELM model. The number of hidden nodes was changed within two intervals, and the runoff simulation effect under different hidden node numbers was analyzed. The empirical formula is expressed as follows:
[0071]
[0072] Where γ is the number of neurons in the hidden layer, s is the number of runoff samples, n is the number of input layers, m is the number of output layer categories, and k is a constant ranging from [2, 10].
[0073] As a preferred technical solution, step S4 is specifically as follows:
[0074] S41. Migration application objects include reservoirs with similar or different hydrogeological conditions in the same area as the study reservoir, as well as hydrological station control sections;
[0075] S42. The effect of model transplantation application is evaluated using four indicators: root mean square error, Nash efficiency coefficient, true positive rate, and false positive rate. The details are as follows:
[0076] The root mean square error represents the average deviation between simulation and observation, with lower values indicating better model performance;
[0077] The Nash efficiency coefficient ranges between 1 and 0. According to the Pearl River Resources Committee of the Ministry of Water Resources, the national hydrological forecast standard requires that the NSE be greater than 0.7, which is considered "acceptable";
[0078] True positive rate
[0079] Where TP is the number of true positives, indicating that both the forecast runoff and the observed value exceeded the threshold. This is a key number used to distinguish "flood" instances. FN is the number of false negatives, indicating that the simulated operational event was not actually a "flood" event. A higher TPR can reduce operational risk.
[0080] False positive rate
[0081] Where FP is the number of false positives, or “false alarms,” indicating that a flood was predicted when there was none, and TN is the number of true negatives, indicating that neither the forecast nor the observations predicted a flood; a lower FPR will increase the benefit of the reservoir because higher hydraulic heads contribute to hydropower generation.
[0082] In a second aspect, the present invention provides a reservoir inflow runoff forecasting system based on machine learning, which is applied to the reservoir inflow runoff forecasting method based on machine learning, and includes a data acquisition module, a MOELM model construction module, a sensitivity analysis module, a model transplantation module, and a runoff forecasting module;
[0083] The data acquisition module is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set;
[0084] The MOELM model construction module is used to construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II) to optimize the runoff simulation effect with the goal of minimizing the root mean square error of runoff overall fitting and extreme flood events. Specifically, it includes:
[0085] Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated;
[0086] The ELM model is constructed with the initially generated IW and B as parameters, and the simulated runoff is obtained with the historical inflow runoff series as input;
[0087] The relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated based on the simulated runoff and the measured runoff;
[0088] Based on the NSGA-Ⅱ algorithm, a multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0089] The IW and B obtained by the last iteration optimization of NSGA-II and the simulated runoff obtained according to the ELM model are output to obtain the trained MOELM model;
[0090] The sensitivity analysis module is used to perform sensitivity analysis on the MOELM model from the perspectives of model input information and the number of hidden nodes by using the partial mutual information method PMI and the empirical guide;
[0091] The model transplantation module is used to migrate the trained MOELM model to other reservoirs to test the portability of the model;
[0092] The runoff forecast module is used to forecast reservoir inflow runoff based on the MOELM model with sensitivity analysis and portability test.
[0093] In a third aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the reservoir inflow runoff forecasting method based on machine learning.
[0094] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0095] The present invention introduces a hybrid machine learning method, MOELM, in the field of hydrological forecasting. It uses the extreme learning machine (ELM) as the framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-Ⅱ). It optimizes the runoff simulation effect with the goals of overall runoff fitting and minimizing the root mean square error of extreme flood events. The model is analyzed from the perspectives of model input information and the number of hidden nodes through the partial mutual information (PMI) method.
[0096] Compared with traditional hydrological models, this method reduces operational risks and improves power generation efficiency by lowering the maximum outflow and water level in typical flood events. It also has strong hydrological migration capabilities and is more scalable and robust in flood forecasting. It can effectively improve the accuracy of flood forecasts and reduce simulation errors of extreme floods. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0098] Figure 1 This is a flow chart of a reservoir inflow runoff forecasting method based on machine learning proposed by the present invention;
[0099] Figure 2 This is a comparison chart of the overall performance of MOELM and optimized learning machine OELM in flood event forecasting in the test set with the same lag time;
[0100] Figure 3 There are four typical flood events to verify the MOELM performance result diagram;
[0101] Figure 4 This is a comparison chart of the next-day runoff forecast results of the present invention and the Xin'an River hydrological forecast model;
[0102] Figure 5 This is the comparison result of MOELM after this training applied to other reservoirs and main sections;
[0103] Figure 6 This is a schematic diagram of the structure of a reservoir inflow runoff forecasting system based on machine learning according to an embodiment of the present invention;
[0104] Figure 7 Schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present invention. DETAILED DESCRIPTION
[0105] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0106] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0107] See also Figure 1 This embodiment provides a reservoir inflow runoff forecasting method based on machine learning, comprising the following steps:
[0108] S1. Initialize data and parameters.
[0109] Specifically, reservoir runoff data and related parameters were collected, including the operational time series from January 1, 2008 to October 31, 2023, which included records of 632 flood events (flow rates exceeding 4000m 3 / s), considering the lack of meteorological and hydrogeological data for the analyzed reservoirs, the Markov method was used to enrich the inflow runoff sample series and divide the runoff samples based on the actual situation.
[0110] Furthermore, step S1 specifically includes:
[0111] S11. Collect six-hourly inflow data of the reservoir;
[0112] S12. To enrich the data series, the sample set is enriched based on the Markov chain. The Markov method is used to expand the runoff samples entering the reservoir as follows:
[0113] In the Markov method, the next state of the current state of the sequence depends only on the current state, not on the previous state. The formula is:
[0114] P(Y n+1 |Y1,Y2,...,Y n )=P(Y n+1 |Y n )
[0115] Where Y1, Y2, ... are the state sequences in the Markov method. The equation shows that the state of the sequence at time n+1 is only related to time n.
[0116] The following steps can be taken to expand the runoff sample:
[0117] First, the mean x and standard deviation S of the actual runoff samples are calculated. As the upper and lower boundaries, the runoff samples are divided into five groups from small to large: dry year, relatively dry year, normal year, relatively wet year, and wet year runoff.
[0118] Secondly, the state transition probability matrix of runoff samples of different categories is established to analyze the state transition probability, which is expressed as follows:
[0119]
[0120] In the formula represents the state transition probability of the kth group of runoff samples from state i to state j; Represents the corresponding state transition probability matrix.
[0121] Furthermore, the Markov chain of the discrete sequence can be expressed as x 2 Statistics to test, suppose the sequence under study contains m possible states, use (f i,j )i, j∈E is recorded as the transfer frequency probability matrix, and the sum of each column of the transfer frequency matrix is divided by the sum of each row and column to obtain the "marginal probability" P .j , when m is large, analyze x with "marginal probability" 2 Statistics, at this time x 2 The statistic is subject to degrees of freedom (m-1) 2 x 2 Distribution, given the confidence α, we can get x by looking up the table 2 >x α 2 ((m-1) 2 ) value, when x 2 >x α 2 ((m-1) 2 )(i.e. x 2 The statistic is greater than the degrees of freedom (m-1) 2 The corresponding χ 2 The value is the value of 2 distribution, with a given significance level (confidence) α and degrees of freedom (m-1) 2 When the corresponding critical value is reached, the current runoff sample is considered to meet the “Markov property”, which is expressed as follows:
[0122]
[0123] Then, calculate the autocorrelation coefficient r of each order k , solve the weight ω of the Markov chain with various step sizes k , expressed as follows:
[0124]
[0125] Finally, the required runoff samples are predicted. The weighted sum of the prediction probabilities of the same state is used as the prediction probability P of the index value in the i-th state. i , the predicted probability is expressed as follows:
[0126]
[0127] After determining the index value, add it to the original sequence, that is, realize the runoff sample supplement of the i-th state, and repeat the above steps to supplement the remaining required samples.
[0128] S13. According to the best practice results, the runoff samples were divided into model training set, validation set and test set at a ratio of 70%, 15% and 15% respectively, and input into MOELM model training.
[0129] S2. Construct and train a multi-objective optimized extreme learning machine (MOELM) model for reservoir inflow forecasting. The MOELM model uses an extreme learning machine (ELM) as its framework and combines it with a fast elite multi-objective genetic algorithm (NSGA-II). It optimizes runoff simulation by minimizing the root mean square error (RMS) of runoff overall fit and extreme flood events. Specifically, the model includes:
[0130] S21. Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated;
[0131] Furthermore, the specific steps of initializing the population of the NSGA-Ⅱ multi-objective optimization algorithm are as follows:
[0132] Initialize the population: Randomly generate a certain number of individuals as the initial population. Each individual represents a potential solution, which can be expressed as follows:
[0133]
[0134] Where x is a single individual in NSGA-Ⅱ, x is the upper and lower bounds of the individual value range, which in this invention refers to the connection weight IW and the bias value B, and is expressed as follows:
[0135]
[0136] where ω n,1 is the nth input layer, the weight of the first hidden layer in the input weight, b n,1 is the bias value of the nth input layer and the first hidden layer.
[0137] S22, using the initially generated IW and B as parameters, construct the ELM model, and use the historical inflow runoff series as input to obtain the simulated runoff;
[0138] Furthermore, step S22 is specifically as follows:
[0139] According to the randomly generated initialization weights IW and bias B, the input layer data is mapped from its original space to the feature space of the hidden layer through the activation function, and then the output weight OW is determined according to the formula, and finally the output layer output ELM(X; L) is obtained, as follows:
[0140] ELM(X;L)=OW×g(IW×X+B)
[0141] ELM(X;L) represents simulated runoff, Q t , X=[Q t-1 ,Q t-2 ,...,Q t-k ] is the input variable, i.e., the historical inflow runoff; IW, B, OW, and g(·) represent the input weight, bias, output weight, and activation function, respectively; and L represents the number of hidden neurons. According to ELM theory, the model ideally approximates all training samples with non-zero error, meaning that the simulated output matches the expected output. Its mathematical expression is:
[0142] ELM(X;L)=Q t
[0143] From the perspective of matrices, the above formula can be written as:
[0144] Q=Oβ
[0145] in:
[0146]
[0147] The value of the output weight OW is uniquely determined using the formula (i.e. ), as follows:
[0148]
[0149] in The Moore-Penrose generalized inverse matrix of β is used to train the parameters IW, B, and To predict future runoff.
[0150] S23, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series calculated based on the simulated runoff and the measured runoff;
[0151] Furthermore, the simulation errors of the entire runoff series and the extreme flood runoff series are expressed as follows:
[0152]
[0153] in and The i th The observed and simulated runoff is n observations, where n is the total number of observations in the validation set. threshold is the threshold for determining extreme flood events. When the runoff data is greater than or equal to it, the flood is considered to be an extreme flood event. The Q threshold 4000m 3 RMSE1 and RMSE2 represent the root mean square error (RMSE) of total runoff and extreme flood events, respectively, that is, the simulation errors of the entire runoff series and the extreme flood runoff series.
[0154] S24. Based on the NSGA-Ⅱ algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series;
[0155] Furthermore, the specific method of multi-objective optimization is as follows:
[0156] Non-dominated sorting: Based on the ELM model, the runoff simulation aims to minimize the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. The population is quickly sorted according to the non-dominated relationship and the population is divided into different levels. The goal is expressed as follows:
[0157] obj1=minRMSE1
[0158] obj2=minRMSE2
[0159] Elite strategy: The best solution of the previous generation is retained as part of the next generation to expand the sampling space and prevent the loss of the best individuals.
[0160] Selection operation: Perform selection operations based on the fitness and non-dominated hierarchy of individuals to retain excellent individuals.
[0161] Mutation operation: After the initial population is generated, the mutation operation is performed, that is, the IW and B value selection parts in each target are mutated. For each target x i,G (i=1,2,...,NP), the mutation vector is generated as follows:
[0162]
[0163] Among them, the randomly selected individual numbers r1, r2, r3 are required to be different from each other and also different from the target vector number i, so the population size must satisfy NP≥4, and the mutation operator F∈[0,2] is a real constant factor that controls the scaling of the deviation variable.
[0164] From the perspective of runoff simulation, it can also be expressed as:
[0165]
[0166] Crossover operation: Cross the mutated population with the initial population to generate a new generation of population, as follows:
[0167]
[0168] Among them, rand(0,1) represents a random value between [0,1], and CR represents the crossover operator, whose value range is [0,1].
[0169] From the perspective of runoff simulation, it can also be expressed as:
[0170]
[0171] Crowding comparison: Compare the crowding levels of individuals in adjacent levels and determine the survival ability of individuals based on the crowding levels.
[0172] Iterative optimization: Repeat the above steps until the termination condition is met (such as reaching the maximum number of iterations or converging to a satisfactory Pareto front solution)
[0173] S25, outputting IW and B obtained by the last iterative optimization of NSGA-II and the simulated runoff obtained according to the ELM model to obtain a trained MOELM model;
[0174] S3, sensitivity analysis of the MOELM model was performed from the perspectives of model input information and number of hidden nodes using partial mutual information method (PMI) and empirical guidelines;
[0175] Furthermore, step S3 is specifically as follows:
[0176] S31. For sensitivity analysis, the partial mutual information method PMI was selected, which is expressed as follows:
[0177] PMI(X,Y|Z)=D(p(x,y,z)||p * (x|z)p * (y|z)p(z))
[0178] Where p(x,y,z) is the joint probability distribution of random variables X,Y,Z, and D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) means p(x,y,z) to p * (x|z)p * The distance between (y|z)p(z).
[0179] In the runoff forecast simulation, for the sensitivity analysis of the input layer structure, the partial mutual information method is first used to preliminarily screen out the set of predictors with the strongest correlation with runoff from the set of predictors. Then, based on the predictor combinations obtained from the initial screening, the correlation coefficient method is used to calculate the correlation coefficients between the predictors and runoff at different lags to identify the most important input predictor set. The sensitivity is verified by analyzing the correlation coefficients and input importance of the selected factor set in the random forest (RF) model.
[0180] S32. For the sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the hidden layer experience guide, the input selected by PMI was used to test the number of hidden layer nodes in the MOELM model. The number of hidden nodes was changed within two intervals, and the runoff simulation effect under different hidden node numbers was analyzed. The empirical formula is expressed as follows:
[0181]
[0182] Where γ is the number of neurons in the hidden layer, s is the number of runoff samples, n is the number of input layers, m is the number of output layer categories, and k is a constant ranging from [2, 10].
[0183] The maximum PMI value (mutual information from the initial round) and 95% confidence limits for each round are listed in Table 1. The correlation coefficients and input importance rankings of the selected inputs with the random forest (RF) model are very good, indicating that these inputs are indeed significant.
[0184] However, the selected candidate inputs have considerable overlap with those in the lag 1-4 scenario, which may lead to slight differences in the evaluation metrics. To address this issue, a lag 4-7 scenario is proposed to demonstrate the performance of four equivalent inputs. Results show that the inputs selected by PMI consistently achieve the best performance across all four input scenarios. Specifically, compared to the lag 4-7 scenario, the overall RMSE1, RMSE2, NSE, TPR, and FPR improve by 12.7%, 35.5%, 8.9%, 29.2%, and 32.2%, respectively, and compared to the lag 1-4 scenario by 1.1%, 1.4%, 5.6%, 2.8%, and 13.2%, respectively. These results confirm the intuition behind the selection of input variables. While the differences between the PMI and lag 1-4 scenarios are not substantial, the selected input variables generally align with the most correlated input variables in the series.
[0185] Table 1 Stepwise predictor selection with partial mutual information (PMI)
[0186]
[0187] S4, migrate the trained MOELM model to other reservoirs to test the portability of the model;
[0188] Furthermore, the portability of the test model is specifically as follows:
[0189] The application objects of S8-1 migration include reservoirs with similar or different hydrogeological conditions in the same area as the study reservoir, as well as hydrological station control sections.
[0190] The effect of S8-2 model transplantation application was evaluated using four indicators: root mean square error, Nash efficiency coefficient, and true and false positive rates. The specific indicators are as follows:
[0191] Root mean square error (RMSE)
[0192]
[0193] RMSE represents the average deviation between simulation and observation, and the lower the value, the better the performance. is the simulated runoff volume in the ith period, is the observed runoff in the i-th period.
[0194] Nash-Sutcliffe efficiency coefficient (NSE)
[0195]
[0196] in is the mean of the observed values. NSE ranges between 1 (perfect fit) and 0. According to the regulations of the Pearl River Resources Committee of the Ministry of Water Resources, the national hydrological forecast standard requires NSE to be greater than 0.7, which is considered "acceptable."
[0197] True Positive Rate (TPR)
[0198]
[0199] Where TP is the number of true positives, indicating that both the forecast runoff and the observed value exceeded the threshold, a key number used to distinguish “flood” instances. FN is the number of false negatives, indicating that the simulated operational event was not actually a “flood” event. A higher TPR can reduce operational risk, as unexpected flooding can burden operational plans (Huang et al. 2022).
[0200] False positive rate (FPR)
[0201]
[0202] Where FP is the number of false positives, or "false alarms," indicating that a flood was predicted when none actually occurred. TN is the number of true negatives, indicating that neither the forecast nor the observations predicted a flood. A lower FPR increases the benefit of a reservoir because higher hydraulic heads contribute to hydropower generation.
[0203] S5. The MOELM model is used to predict reservoir inflow based on sensitivity analysis and portability test.
[0204] Comparing the optimized extreme learning machine OELM's simulation forecast results of the reservoir flood event and the differences with the actual flood, the results are as follows Figure 2 As shown, the prediction accuracy obtained by MOELM exceeds that of OELM. In the scenarios with different input delay times, the improvement is about 5.27% on average. The improvement is not obvious at first, but it continues to improve, highlighting the importance of input selection. Although MOELM slightly increases the RMSE of all records, the margin is negligible at around 0.9%, and the NSE value only slightly decreases from 0.8 to 0.79. This difference is acceptable for subsequent reservoir operation under extreme events; the four types of measured flood events selected are divided into the following: two types of flood peak flows are about 10,000m 3 / s, representing a 20-year flood period, with two peak flows of approximately 6000m 3 / s, representing a 10-year flood period, has become a reliable flood forecast that is consistent with observations. The deviation of the peak flow is less than 5%.
[0205] like Figure 3As shown, MOELM generally achieved the best performance across all tests, comparing the simulation results of linear regression (LR), single-layer feedforward network (SLFN), extreme gradient boosting (XGboost), and random forest (RF). Table 2 lists the averages of all 20 runs with different lag times. The results show that the algorithms based on RMSE, NSE, and TPR are significantly different based on the group t-test results with p=0.05. However, the FPR shows an opposite trend, that is, the best algorithm does not achieve the smallest FPR. The false alarm rate in the MOELM model is quite low, approximately 0.2%.
[0206] Table 2 Average running time of different methods
[0207]
[0208]
[0209]
[0210] Comparing the simulation results of the Xinanjiang model, reviewing the observed historical runoff, analyzing the results of maximum water level and maximum reservoir output, and hydropower operation risk, the results are as follows: Figure 4 As shown, the maximum flow rate in the relevant scenario calculations increased by an average of 490m 3 / s, with an average water level increase of 1.3m. In comparison, the risks associated with MOELM are negligible. For hydropower generation, using MOELM's predicted flow, hydropower stations can generate an additional 130 million kilowatt-hours.
[0211] In order to evaluate the transferability of the model, the present invention was applied to the three large reservoirs of Tianyi, Baise and Changzhou in the Pearl River Basin and the control section of Wuzhou Station. The results are as follows: Figure 5 As shown, the model demonstrates excellent transferability. The NSE values for Tianyi and Baise Reservoirs are 0.75 and 0.65, respectively, which are 6.5% and 18.7% lower than those for Longtan Reservoir. Tianyi Reservoir has similar meteorological and hydrological conditions to Longtan Reservoir, resulting in less variation. Conditions at Baise Reservoir are significantly different, resulting in a significant decrease in model performance. The model performs very well in Wuzhou and Changzhou, with NSE values exceeding 0.95. This superior performance is attributed to changes in flow direction. The rates of change (relative to the average level) for Wuzhou and Changzhou are approximately 8.5% and 4.7%, respectively, both significantly lower than the 56.3% for Longtan Reservoir.
[0212] The present invention can accurately predict the next day's runoff of a reservoir, and provide a more accurate method for runoff forecast analysis.
[0213] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0214] Based on the same concept as the machine learning-based reservoir inflow runoff forecasting method in the above-mentioned embodiment, the present invention also provides a machine learning-based reservoir inflow runoff forecasting system, which can be used to execute the above-mentioned machine learning-based reservoir inflow runoff forecasting method. For ease of explanation, the structural schematic diagram of the embodiment of the machine learning-based reservoir inflow runoff forecasting system only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and it can include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0215] See also Figure 6 In another embodiment of the present application, a reservoir inflow runoff forecasting system 100 based on machine learning is provided, the system comprising a data acquisition module 101, a MOELM model construction module 102, a sensitivity analysis module 103, a model transplantation module 104 and a runoff forecasting module 105;
[0216] The data acquisition module 101 is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set;
[0217] The MOELM model construction module 102 is used to construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as a framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II) to optimize the runoff simulation effect with the goal of minimizing the root mean square error of runoff overall fitting and extreme flood events. Specifically, it includes:
[0218] Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated;
[0219] The ELM model is constructed with the initially generated IW and B as parameters, and the simulated runoff is obtained with the historical inflow runoff series as input;
[0220] The relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated based on the simulated runoff and the measured runoff;
[0221] Based on the NSGA-Ⅱ algorithm, a multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series.
[0222] The IW and B obtained by the last iteration optimization of NSGA-II and the simulated runoff obtained according to the ELM model are output to obtain the trained MOELM model;
[0223] The sensitivity analysis module 103 is used to perform sensitivity analysis on the MOELM model from the perspectives of model input information and the number of hidden nodes by using the partial mutual information method PMI and the empirical guide;
[0224] The model transplantation module 104 is used to migrate the trained MOELM model to other reservoirs to test the portability of the model;
[0225] The runoff forecasting module 105 is used to forecast the reservoir inflow runoff based on the MOELM model with sensitivity analysis and portability test.
[0226] It should be noted that the reservoir inflow runoff forecasting system based on machine learning of the present invention corresponds one-to-one to the reservoir inflow runoff forecasting method based on machine learning of the present invention. The technical features and beneficial effects described in the embodiment of the above-mentioned reservoir inflow runoff forecasting method based on machine learning are all applicable to the embodiment of the reservoir inflow runoff forecasting method based on machine learning. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.
[0227] In addition, in the implementation of the machine learning-based reservoir inflow runoff forecasting system in the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the machine learning-based reservoir inflow runoff forecasting system is divided into different program modules to complete all or part of the functions described above.
[0228] See also Figure 7 In one embodiment, this embodiment provides a computer-readable storage medium storing a program. When the program is executed by a processor, the method for predicting reservoir inflow based on machine learning is implemented, specifically:
[0229] S1. Collect reservoir runoff data and expand the runoff data to obtain a runoff sample set;
[0230] S2. Construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II). It optimizes the runoff simulation effect by minimizing the root mean square error of the overall runoff fit and extreme flood events. Specifically, it includes:
[0231] S21. Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated;
[0232] S22, using the initially generated IW and B as parameters, construct the ELM model, and use the historical inflow runoff series as input to obtain the simulated runoff;
[0233] S23, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series calculated based on the simulated runoff and the measured runoff;
[0234] S24. Based on the NSGA-Ⅱ algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series;
[0235] S25, outputting IW and B obtained by the last iterative optimization of NSGA-II and the simulated runoff obtained according to the ELM model to obtain a trained MOELM model;
[0236] S3, sensitivity analysis of the MOELM model was performed from the perspectives of model input information and number of hidden nodes using partial mutual information method (PMI) and empirical guidelines;
[0237] S4, migrate the trained MOELM model to other reservoirs to test the portability of the model;
[0238] S5. The MOELM model is used to predict reservoir inflow based on sensitivity analysis and portability test.
[0239] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0240] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0241] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A reservoir inflow forecasting method based on machine learning, characterized in that: The steps include: S1. Collect reservoir runoff data and expand the runoff data to obtain a runoff sample set; S2. Construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II). It optimizes the runoff simulation effect by minimizing the root mean square error of the overall runoff fit and extreme flood events. Specifically, it includes: S21. Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated; S22, using the initially generated IW and B as parameters, construct the ELM model, and use the historical inflow runoff series as input to obtain the simulated runoff; S23, the relative error of the entire runoff series and the simulation error of the extreme flood runoff series calculated based on the simulated runoff and the measured runoff; S24. Based on the NSGA-Ⅱ algorithm, multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series; S25, outputting IW and B obtained by the last iterative optimization of NSGA-II and the simulated runoff obtained according to the ELM model to obtain a trained MOELM model; S3, sensitivity analysis of the MOELM model was performed from the perspectives of model input information and number of hidden nodes using partial mutual information method (PMI) and empirical guidelines; S4, migrate the trained MOELM model to other reservoirs to test the portability of the model; S5. The MOELM model is used to predict reservoir inflow based on sensitivity analysis and portability test.
2. The reservoir inflow forecasting method based on machine learning according to claim 1 is characterized in that: In step S1, the runoff sample set is expanded based on Markov, specifically: Statistical mean of actual runoff samples and standard deviation S, with As the upper and lower bounds, the runoff samples are divided into five groups from small to large: dry year, relatively dry year, normal year, relatively wet year, and wet year runoff; Establish the state transition probability matrix of different types of runoff samples and analyze the state transition probability; The Markov chain of discrete sequences can be expressed as x 2 Statistics to test, suppose the sequence under study contains m possible states, use (f i,j )i is recorded as the transfer frequency probability matrix, j∈E, and the sum of each column of the transfer frequency matrix is divided by the sum of each row and column to obtain the "marginal probability" P .j , when m is large, analyze x with "marginal probability" 2 Statistics, at this time x 2 The statistic is subject to degrees of freedom (m-1) 2 x 2 Distribution, given the confidence α, we can get x by looking up the table 2 >x α 2 ((m-1) 2 ) value, when x 2 >x α 2 ((m-1) 2 ) represents x 2 The statistic is greater than the degrees of freedom (m-1) 2 The corresponding χ 2 The value refers to the 2 distribution, with a given confidence level α and degrees of freedom (m-1) 2 When the corresponding critical value is reached, the current runoff sample is considered to meet the "Markov property"; Calculate the autocorrelation coefficients of each order and solve the weights of Markov chains with various step sizes; Predict the required runoff samples, and take the weighted sum of the prediction probabilities of the same state as the prediction probability P of the index value in the i-th state i ; After determining the index value, add the index value to the original sequence, that is, realize the runoff sample supplement of the i-th state, repeat the above steps to supplement the remaining required samples, and finally obtain the runoff sample set; The runoff samples in the runoff sample set are divided into training set, validation set and test set according to the set ratio, which serve as the input for MOELM model training.
3. The reservoir inflow forecasting method based on machine learning according to claim 1 is characterized in that: Step S21 also includes the step of initializing the population using the NSGA-II multi-objective optimization algorithm, specifically: Initialize the population: Randomly generate a certain number of individuals as the initial population. Each individual represents a potential solution, which can be expressed as follows: Where x is a single individual in NSGA-Ⅱ, x The upper and lower bounds of the individual value range refer to the connection weight IW and the bias value B, which are expressed as follows: where ω n,1 is the nth input layer, the weight of the first hidden layer in the input weight, b n,1 is the bias value of the nth input layer and the first hidden layer.
4. The reservoir inflow forecasting method based on machine learning according to claim 1 is characterized in that: The step S22 is specifically as follows: According to the randomly generated initialization weights IW and bias B, the input layer data is mapped from its original space to the feature space of the hidden layer through the activation function, and then the output weight OW is determined according to the formula, and finally the output layer output ELM(X; L) is obtained, as follows: ELM(X;L)=OW×g(IW×X+B) ELM(X;L) represents simulated runoff; Q t , X=[Q t-1 ,Q t-2 ,...,Q t-k ] is the input variable, representing the historical inflow runoff; IW, B, OW, and g(·) represent the input weight, bias, output weight, and activation function, respectively; L represents the number of hidden neurons; According to the ELM theory, the simulated output matches the expected output, and the mathematical expression is: ELM(X;L)=Q t From the perspective of the matrix, the above formula can be rewritten as: Q=Oβ in: Output weight OW value It is uniquely determined using the formula, as follows: in The Moore-Penrose generalized inverse matrix of β is used to train the parameters IW, B, and To predict future runoff.
5. The reservoir inflow forecasting method based on machine learning according to claim 1 is characterized in that: In step S23, the simulation errors of the entire runoff series and the extreme flood runoff series are expressed as follows: in, and The i th The observed and simulated runoff is n observations, n is the total number of observations in the validation set; Q threshold is the threshold for determining extreme flood events. When the runoff data is greater than or equal to Q threshold When , this flood is considered to be an extreme flood event. RMSE1 and RMSE2 represent the root mean square error RMSE of the total runoff and the extreme flood event, respectively, that is, the simulation errors of the entire runoff series and the extreme flood runoff series.
6. The reservoir inflow forecasting method based on machine learning according to claim 5 is characterized in that: In step S24, the multi-objective optimization is specifically performed as follows: Non-dominated sorting: Based on the ELM model, the runoff simulation aims to minimize the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. The population is quickly sorted according to the non-dominated relationship and the population is divided into different levels. The goal is expressed as follows: obj1=minRMSE1 obj2=minRMSE2 Elite strategy: retain the best solution of the previous generation as part of the next generation to expand the sampling space and prevent the loss of the best individuals; Selection operation: Perform selection operations based on individual fitness and non-dominated hierarchy to retain excellent individuals; Mutation operation: After the initial population is generated, the mutation operation is performed, that is, the IW and B value selection parts in each target are mutated. For each target x i,G , i=1,2,...,NP, mutation vector v i,G+1 The generation method is as follows: Among them, the randomly selected individual serial numbers r1, r2, r3 are required to be different from each other and also different from the target vector serial number i, so the population size must satisfy NP ≥ 4, and the mutation operator F∈[0,2] is a real constant factor that controls the scaling of the deviation variable; From the perspective of runoff simulation, it can be expressed as: IW and B are input weight and bias respectively; Crossover operation: Cross the mutated population with the initial population to generate a new generation of population ui j,G+1 , as follows: Among them, rand(0,1) represents a random value between [0,1], CR represents the crossover operator, and its value range is [0,1]; From the perspective of runoff simulation, it can be expressed as: Crowding comparison: Compare the crowding levels of individuals in adjacent levels and determine the survival ability of individuals based on the crowding level; Iterative optimization: Repeat the above steps until the termination condition is met.
7. The reservoir inflow forecasting method based on machine learning according to claim 5 is characterized in that: In step S3, sensitivity analysis of the MOELM model is performed from the perspectives of model input information and the number of hidden nodes using the partial mutual information method (PMI) and empirical guidance, specifically: S31. For sensitivity analysis, the partial mutual information method PMI is selected, which is expressed as follows: PMI(X,Y|Z)=D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) Where p(x,y,z) is the joint probability distribution of random variables X,Y,Z, and D(p(x,y,z)||p * (x|z)p * (y|z)p(z)) means p(x,y,z) to p * (x|z)p * The distance of (y|z)p(z); In runoff forecast simulation, for the sensitivity analysis of the input layer structure, the partial mutual information method is first used to preliminarily screen out the set of predictors with the strongest correlation with runoff from the set of predictors. Then, based on the predictor combinations obtained from the initial screening, the correlation coefficient method is used to calculate the correlation coefficients between the predictors and runoff at different lags to identify the most important input predictor sets. The sensitivity is verified by analyzing the correlation coefficients and input importance of the selected factor sets in the random forest RF model. S32. For the sensitivity analysis of the number of hidden layer nodes, based on the empirical formula in the hidden layer experience guide, the input selected by PMI was used to test the number of hidden layer nodes of the MOELM model. The number of hidden nodes was changed within two intervals, and the runoff simulation effect under different hidden node numbers was analyzed. The empirical formula is expressed as follows: Where γ is the number of neurons in the hidden layer, s is the number of runoff samples, n is the number of input layers, m is the number of output layer categories, and k is a constant ranging from [2, 10].
8. The reservoir inflow forecasting method based on machine learning according to claim 1 is characterized in that: The step S4 is specifically as follows: S41. Migration application objects include reservoirs with similar or different hydrogeological conditions in the same area as the study reservoir, as well as hydrological station control sections; S42. The effect of model transplantation application is evaluated using four indicators: root mean square error, Nash efficiency coefficient, true positive rate, and false positive rate. The details are as follows: The root mean square error represents the average deviation between simulation and observation, with lower values indicating better model performance; The Nash efficiency coefficient ranges between 1 and 0. According to the Pearl River Resources Committee of the Ministry of Water Resources, the national hydrological forecast standard requires that the NSE be greater than 0.7, which is considered "acceptable"; True positive rate Where TP is the number of true positives, indicating that both the forecast runoff and the observed value exceeded the threshold. This is a key number used to distinguish "flood" instances. FN is the number of false negatives, indicating that the simulated operational event was not actually a "flood" event. A higher TPR can reduce operational risk. False positive rate Where FP is the number of false positives, or "false alarms," indicating that a flood was predicted when there was none, and TN is the number of true negatives, indicating that neither the forecast nor the observations predicted a flood; a lower FPR will increase the benefit of the reservoir because higher hydraulic heads contribute to hydropower generation.
9. The reservoir inflow forecasting system based on machine learning is characterized by: A reservoir inflow runoff forecasting method based on machine learning applied to any one of claims 1-7, comprising a data acquisition module, a MOELM model construction module, a sensitivity analysis module, a model transplantation module, and a runoff forecasting module; The data acquisition module is used to collect reservoir inflow runoff data and expand the runoff data to obtain a runoff sample set; The MOELM model construction module is used to construct and train a multi-objective optimization extreme learning machine model (MOELM) for reservoir inflow forecasting. The MOELM model uses the extreme learning machine (ELM) as its framework and combines it with the fast elite multi-objective genetic algorithm (NSGA-II) to optimize the runoff simulation effect with the goal of minimizing the root mean square error of runoff overall fitting and extreme flood events. Specifically, it includes: Using the NSGA-Ⅱ multi-objective optimization algorithm as the framework, the connection weights IW and bias values B between the ELM input layer and the hidden layer are randomly generated; The ELM model is constructed with the initially generated IW and B as parameters, and the simulated runoff is obtained with the historical inflow runoff series as input; The relative error of the entire runoff series and the simulation error of the extreme flood runoff series are calculated based on the simulated runoff and the measured runoff; Based on the NSGA-Ⅱ algorithm, a multi-objective optimization is performed with the goal of minimizing the relative error of the entire runoff series and the simulation error of the extreme flood runoff series. The IW and B obtained by the last iteration optimization of NSGA-II and the simulated runoff obtained according to the ELM model are output to obtain the trained MOELM model; The sensitivity analysis module is used to perform sensitivity analysis on the MOELM model from the perspectives of model input information and the number of hidden nodes by using the partial mutual information method PMI and the empirical guide; The model transplantation module is used to migrate the trained MOELM model to other reservoirs to test the portability of the model; The runoff forecast module is used to forecast reservoir inflow runoff based on the MOELM model with sensitivity analysis and portability test.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the reservoir inflow runoff forecasting method based on machine learning described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Flood control optimization scheduling method based on flood type self-identification strategy
CN116502773A
Port rainfall runoff simulation and waterlogging forecasting method and system based on SWMM
CN119294297A