Method for joint dispatching of power plant and network river based on machine learning algorithm and rain flood numerical model
By combining machine learning algorithms with rainfall and flood numerical models, a joint scheduling method for power plants, networks, and rivers was constructed. This method solves the problems of lag and low computational efficiency in traditional scheduling methods, enabling rapid response and optimized scheduling for short-duration extreme rainfall, and reducing losses from urban flooding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE
- Filing Date
- 2022-12-05
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional urban water quantity and quality management lacks a holistic perspective, relies on manual experience, and is difficult to effectively cope with floods caused by short-duration extreme heavy rainfall. Existing rainwater and flood numerical models have low computational efficiency, and optimization algorithms increase costs and lack timeliness.
By combining machine learning algorithms with rainfall and flood numerical models, a joint scheduling method for power plants, networks, and rivers is established. By constructing a rainfall scenario database and a benefit index learning model, the scheduling scheme is optimized to achieve rapid adaptation.
It improves the timeliness and intelligence of joint scheduling of power plants, networks and rivers, assists managers in making quick decisions, and reduces losses from urban flooding.
Smart Images

Figure CN116011731B_ABST
Abstract
Description
A Joint Scheduling Method for Power Plants, Networks, and Rivers Based on Machine Learning Algorithms and Rainfall Numerical Models Technical Field
[0001] This invention relates to the field of flood control and drainage, specifically a method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models. Background Technology
[0002] Global climate change has accelerated the hydrological cycle, leading to more frequent short-duration extreme rainfall events and a significant increase in the probability of urban flooding. While my country has largely completed the construction of sewage treatment plants, drainage networks, and urban river and lake systems, these are typically operated and managed independently by different departments, lacking a unified data-driven operational framework. Traditional urban water quantity and quality management relies on the experience of dispatchers, resulting in subjective scheduling based on actual conditions. This approach lacks a holistic perspective and is inherently lagging. Therefore, traditional scheduling methods are insufficient to fully utilize the flood control and drainage capabilities of the plant, network, and river systems.
[0003] In recent years, combining urban flood disaster simulation with stormwater numerical models and optimizing scheduling schemes using strategies such as particle swarm optimization has become the mainstream method in the field of urban water quantity and quality scheduling and management. However, when simulating large and complex pipe networks, stormwater numerical models have low computational efficiency, and the addition of optimization algorithms further increases computational costs. Their timeliness is often not effectively guaranteed, making it difficult to effectively address urban flood disasters induced by short-duration extreme heavy rainfall. Summary of the Invention
[0004] To improve the efficiency of obtaining scheduling schemes, a joint scheduling method for power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models is proposed.
[0005] The technical solution adopted by the present invention to solve the above problems is:
[0006] A joint scheduling method for power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models includes:
[0007] Step 1: Establish a SWMM model of the target area that couples the sewage treatment plant, urban pipe network and urban river and lake system;
[0008] Step 2: Calibrate the SWMM model parameters based on previous monitoring data of the target area;
[0009] Step 3: Pre-formulate a multi-factory-network-river joint dispatching scheme;
[0010] Step 4: Construct an urban stormwater scenario database: Based on the rainstorm parameters of the target area and the Chicago rain pattern, construct short-duration extreme heavy rainfall events with different return periods to build an urban stormwater scenario database;
[0011] Step 5: Construct an evaluation system for the joint dispatching effect of power plants, grids, and rivers;
[0012] Step 6: Construct a scenario database for joint dispatching of urban power plants, networks, and rivers: Combine the various joint dispatching schemes for power plants, networks, and rivers pre-designed in Step 3 with the short-duration extreme heavy rainfall events constructed in Step 4, and conduct simulations based on the calibrated SWMM model to obtain joint dispatching scenario data for various short-duration extreme heavy rainfall events under different joint dispatching schemes; use the evaluation system for the joint dispatching effect of power plants, networks, and rivers constructed in Step 5 to evaluate the combined schemes and form a scenario database for joint dispatching of urban power plants, networks, and rivers.
[0013] Step 7: Construct a learning model of rainfall conditions, joint dispatch scheme, and benefit index. This model takes rainfall conditions and benefit index as input conditions and the optimal joint dispatch scheme as the output objective. The model is trained based on the urban plant-network-river joint dispatch scenario database constructed in Step 6.
[0014] Step 8: Use the trained rainfall conditions-joint scheduling scheme-benefit index learning model to obtain the optimal joint scheduling scheme for the plant, network and river under the current rainfall conditions.
[0015] Furthermore, step 1 specifically includes:
[0016] Step 11: Based on the river and lake system information of the target area, simplify the river channels into connecting open channels and the lakes into regulating reservoirs;
[0017] Step 12: Connect the outlet of the pipeline to the river and lake, and set the outflow conditions according to the water level of the river and lake.
[0018] Furthermore, the outflow condition in step 12 is specifically as follows:
[0019] When the river water level is lower than the outlet elevation, the outlet discharges water in a free outflow manner, and the flow rate can be expressed as:
[0020]
[0021] In the formula, A is the cross-sectional area of the discharge outlet; H0 is the head height; ε is the lateral contraction coefficient; ζ is the local head loss coefficient; and g is the gravitational acceleration.
[0022] When the river water level is higher than the outlet water level, but the pipeline head is higher than the river head, the outlet discharges as a submerged outflow, and its flow rate can be expressed as:
[0023]
[0024] In the formula, A is the cross-sectional area of the outlet; z is the difference between the pipe head and the river head; ε is the lateral contraction coefficient; ζ is the local head loss coefficient; and g is the gravitational acceleration.
[0025] When the river water level is higher than the outlet and the pipeline head is lower than the river water head, the water in the pipeline will not be discharged from the outlet, and at the same time, the river water will not flow back into the pipeline network.
[0026] Furthermore, the short-duration extreme heavy rainfall events in step 4 are specifically 1-hour, 2-hour, and 3-hour short-duration extreme heavy rainfall events.
[0027] Furthermore, in step 5, the evaluation indicators consist of the cumulative overflow of nodes, the cumulative number of overflowing nodes, the cumulative amount of pollutant overflow and discharge, and the cumulative amount of water pumped by the pumping station. The above indicators are comprehensively considered by constructing an objective function. The constraints are set as follows: the maximum power of the pumping station shall not exceed the rated power of the pumping station, the maximum influent flow of the sewage treatment plant shall not exceed its rated flow, the water level of the regulating lakes in the watershed shall not exceed the highest operating water level, and the water level of the flood control and drainage river shall not exceed the maximum safe water level.
[0028] Furthermore, the specific steps of step 7, constructing the learning model of rainfall conditions-joint scheduling scheme-benefit index, are as follows:
[0029] Step 71: Calculate the characteristic parameters of the rainfall time series based on rainfall conditions;
[0030] Step 72: Obtain the optimal scheduling scheme under each rainfall condition from the urban plant-network-river joint scheduling scenario database;
[0031] Step 73: Calculate the correlation between the feature parameters of the rainfall time series and each optimal scheduling scheme, and select the parameters with a correlation coefficient greater than the preset value as the training parameters of the machine learning model;
[0032] Step 74: Divide the training parameters into a training set D and a test set according to a certain ratio;
[0033] Step 75: Establish a machine learning model and train the model using a training set; the machine learning model uses the training set data as input data and the gate opening degree, pump station operating power, and total sewage treatment plant influent in the joint scheduling scheme as target data for model training.
[0034] Furthermore, the machine learning models in step 75 include the K-nearest neighbor model, the random forest model, and the extreme random tree model; step 8 takes the scheduling scheme with the smallest objective function among the three model output schemes as the final optimized scheduling scheme.
[0035] Furthermore, the steps for training the K-nearest neighbor model using the training set are as follows:
[0036] The training samples are represented in the format (x, f(x)), where x is the feature parameter of the sample, and x is derived from (x 1 ,x 2 ,x 3 ,…,x n Composed of ) where x n Let x be the nth attribute value of sample x; for a new input sample x i x is calculated one by one using the Euclidean distance formula. i The distance between x and each sample in the training set is used to select samples that are similar to x. i The K nearest samples; Euclidean distance is expressed as:
[0037]
[0038] In the formula, x i x j There are two samples respectively; Sample x i and x j The l-th eigenvalue; L(x) j ,x j ) is the sample x i and x j The distance between them.
[0039] Furthermore, the steps for training a random forest model using the training set are as follows:
[0040] Multiple parallel training groups are generated using the Bootstrap resampling method, and the decision tree model is trained independently: First, the empirical entropy H(D) of the training set D is calculated:
[0041]
[0042] Calculate the empirical conditional entropy H(D|A) of feature A on training set D:
[0043]
[0044] Calculate information gain:
[0045] g(D,A)=H(D)-H(D|A)
[0046] Calculate the information gain ratio:
[0047]
[0048] in:
[0049]
[0050] In the formula, D represents the entire training dataset, A represents the feature parameters, K represents the total number of categories, and C represents the total number of categories. k For class k, n is the number of values for feature parameter A;
[0051] Select the feature parameter A with the largest information gain ratio. g As a node, for feature parameter A g Possible values {a1, a2, ..., a n}, according to A in sequence g =a1,…,A=a n The training set D is split into Di and Di. i ,D2,…,D n Enter the next level, with A-{A g} represents the feature parameter set. Repeat the above steps until all feature parameters have been traversed, and then output the decision tree model.
[0052] All independently generated decision tree models are combined to construct a random forest model.
[0053] Furthermore, the steps for training the extreme random tree model using the training set are as follows:
[0054] Select all training sets D for model training. Randomly select N feature parameters from feature parameters A, and randomly select one feature parameter as a splitting node. Take the splitting threshold with the smallest Gini coefficient as the optimal splitting threshold. Generate two child nodes from the data D_left and D_right, and iterate through the remaining parameters in turn. The Gini coefficient of the training set D can be expressed as:
[0055]
[0056] The Gini coefficient of set D under characteristic parameter A can be expressed as:
[0057]
[0058] In the formula, D represents the entire training dataset, A represents the feature parameters, K represents the total number of categories, and C represents the total number of categories. k For class k, D1 and D2 are two subsets divided based on feature A.
[0059] The beneficial effects of this invention compared to the prior art are as follows: by simulating the scheduling effects of various joint scheduling methods under various rainstorm conditions, and further combining machine learning algorithms to construct a learning model of rainfall scenario-joint scheduling scheme-benefit index, the joint scheduling scheme can be quickly adapted to rainfall events, improving the timeliness and intelligence of joint scheduling of power plants, networks and rivers, assisting managers to make quick decisions in the short term, and reducing urban flood disaster losses. Attached Figure Description
[0060] Figure 1 is a flowchart of the joint scheduling method of power plants, networks and rivers based on machine learning algorithms and rainfall-flood numerical models;
[0061] Figure 2 is a flowchart of the construction process of the urban plant-network-river joint scheduling scenario database;
[0062] Figure 3 is a flowchart of the learning model construction process for rainfall conditions, joint scheduling schemes, and benefit indicators. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0064] As shown in Figure 1, the joint scheduling method of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models includes:
[0065] Step S1: Collect and organize basic data for the target area, including topography, meteorology, drainage facilities, and river and lake system data, including digital elevation data, land use data, typical rainstorm sequences, drainage pipeline data, drainage node data, pumping station information, sewage treatment plant information, river and lake system distribution information, and water-related engineering facility information.
[0066] Step S2: Collect and organize monitoring data for the target area, including rainfall monitoring data, pipeline flow and water level monitoring data, pipeline outlet flow monitoring data, river water level and flow monitoring data, lake water level monitoring data, pump station monitoring data, and sewage treatment plant monitoring data.
[0067] Step S3: Establish a SWMM model of the target area, coupling the wastewater treatment plant, urban pipe network, and urban river and lake system. The specific steps are as follows:
[0068] Step S3.1: Based on the river and lake system information, simplify the river channels into connecting open channels and the lakes into regulating reservoirs;
[0069] Step S3.2: Connect the outlet of the pipeline to the river or lake, and set the outflow conditions based on the water level of the river or lake. The specific outflow rules are as follows:
[0070] 1) When the river level is lower than the outlet elevation, the discharge method is free outflow, and the flow rate can be expressed as:
[0071]
[0072] In the formula, A is the cross-sectional area of the discharge outlet; H0 is the head height; ε is the lateral contraction coefficient; ζ is the local head loss coefficient; and g is the gravitational acceleration.
[0073] 2) When the river level is higher than the outlet, but the pipeline head is higher than the river head, the outlet discharges as a submerged outflow, and its flow rate can be expressed as:
[0074]
[0075] In the formula, A is the cross-sectional area of the outlet; z is the difference between the pipe head and the river head; ε is the lateral contraction coefficient; ζ is the local head loss coefficient; and g is the gravitational acceleration.
[0076] 3) When the river water level is higher than the outlet and the pipeline head is lower than the river water head, the control gate will be opened by default, meaning that the water in the pipeline will not be discharged from the outlet, and at the same time, the water in the river will not flow back into the pipe network.
[0077] Step S4: Calibrate SWMM model parameters based on previous monitoring data of the target area: Based on the monitoring data of the target area collected in Step S2, using rainfall monitoring data, river water level and flow monitoring data, lake water level monitoring data, pump station monitoring data and sewage treatment plant monitoring data as model driving conditions, and pipeline flow and water level monitoring data and pipeline outlet flow monitoring data as verification information, calibrate the SWMM model parameters to ensure that the established model has high reliability and can accurately invert the process of rainstorm-induced flooding.
[0078] Step S5: Pre-prepared multi-factory network river joint dispatch scheme: Based on the number and location of pumping stations, sewage treatment plants, and sluice gates, multiple pre-prepared dispatch schemes are set up in advance to enable the dispatch schemes to have as many flexible combinations as possible to meet the joint dispatch needs under various rainfall conditions.
[0079] Step S6: Construct an urban stormwater scenario database. Based on the rainstorm parameters of the study area and the Chicago rainfall pattern, construct short-duration extreme heavy rainfall events with different return periods to fully cover all possible short-duration extreme rainfall events in the target area, taking into account various rainfall conditions. In this embodiment, 1-hour, 2-hour, and 3-hour short-duration extreme heavy rainfall events are constructed.
[0080] Step S7: Construct an evaluation system for the joint scheduling effect of the plant, network, and river. In this embodiment, the evaluation indicators for the joint scheduling scheme consist of the cumulative overflow of nodes, the cumulative number of overflow nodes, the cumulative pollutant overflow and discharge, and the cumulative water pumping volume of pumping stations. An objective function is constructed to comprehensively consider these indicators. Constraints are set as follows: the maximum power of the pumping station cannot exceed its rated power; the maximum influent flow of the sewage treatment plant cannot exceed its rated flow; the water level of the regulating lakes within the watershed does not exceed the highest operating water level; and the water level of the flood control and drainage channels does not exceed the maximum safe water level. This allows for a comprehensive evaluation of the scheduling scheme's effectiveness and the selection of the optimal scheduling scheme. Other indicators can also be selected for evaluation according to actual needs; no restrictions are imposed here.
[0081] The specific evaluation indicators for the plant-network-river scheduling scheme are as follows:
[0082] 1) Nodal cumulative overflow: Nodal cumulative overflow is used to measure the amount of surface inundation water, and is expressed by the following formula:
[0083]
[0084] In the formula q it Let Δt be the overflow of the i-th node per unit time at time t, where Δt is the time step, N is the cumulative number of nodes, and T is the cumulative simulation time.
[0085] 2) Cumulative number of overflow nodes: The cumulative number of overflow nodes is used to measure the extent of surface flooding and is expressed by the following formula:
[0086]
[0087] In the formula q it Let Δt be the overflow of the i-th node per unit time at time t, where Δt is the time step, N is the cumulative number of nodes, and T is the cumulative simulation time.
[0088] 3) Cumulative Pollutant Overflow and Discharge: Cumulative pollutant overflow and discharge is used to measure the degree of pollutant control. A high cumulative pollutant overflow and discharge indicates poor pollutant control. It is expressed by the following formula:
[0089]
[0090] In the formula q it Let out be the overflow of the i-th node per unit time at time t. jt Let c be the overflow rate of the j-th discharge outlet per unit time at time t. it and c jtLet be the pollutant concentrations at node i and outlet j at time t, respectively; Δt be the time step; N and M be the cumulative number of nodes and the cumulative number of outlets, respectively; and T be the cumulative simulation time.
[0091] 4) Cumulative water pumping volume of the pumping station: The cumulative water pumping volume of the pumping station is used to account for the main energy consumption during the dispatching process, and is expressed by the following formula:
[0092]
[0093] In the formula q kt Let Δt be the overflow of the i-th node per unit time at time t, where Δt is the time step, N is the cumulative number of nodes, and T is the cumulative simulation time.
[0094] The optimal objective function is set as follows:
[0095]
[0096] in:
[0097] w1 + w2 + w3 + w4 = 1
[0098]
[0099]
[0100]
[0101]
[0102] In the formula These represent the normalized values of the cumulative overflow flow, the cumulative number of overflowing nodes, the cumulative pollutant overflow and discharge, and the cumulative water pumping volume of the pumping station, respectively. w1, w2, w3, and w4 are their respective weighting parameters.
[0103] The constraints, namely that the maximum power of the pumping station cannot exceed the rated power of the pumping station, the maximum influent flow of the sewage treatment plant cannot exceed its rated flow, the water level of the regulating lakes within the watershed cannot exceed the highest operating water level, and the water level of the flood control and drainage channels cannot exceed the maximum safe water level, can be expressed as:
[0104]
[0105] In the formula, Z j Let the water level of the j-th river within the flood control area be denoted as . This is the maximum safe water level determined based on the flood control and drainage requirements of the urban watershed. These are the maximum safe water levels determined based on flood control and drainage requirements.
[0106] Step S8: Construct a scenario database for joint dispatching of urban power plants, networks, and rivers: Combine the pre-designed joint dispatching schemes for power plants, networks, and rivers proposed in Step S5 with the short-duration extreme heavy rainfall events proposed in Step S6. Simulate using the SWMM model calibrated in Step S4 to obtain joint dispatching scenario data for various extreme rainfall events under different dispatching schemes. Calculate the joint dispatching effect evaluation index proposed in Step S7 and organize it to form a scenario database for joint dispatching of urban power plants, networks, and rivers. The flowchart is shown in Figure 2, where N is the number of rainfall events and M is the number of joint dispatching schemes.
[0107] Step S9: Constructing a learning model for rainfall conditions, joint scheduling schemes, and benefit indicators: This embodiment constructs a model based on machine learning algorithms. As shown in Figure 3, three machine learning algorithms—K-nearest neighbors, random forest, and extreme random tree—are used to establish the model between rainfall conditions, joint scheduling schemes, and benefit indicators. The model uses rainfall conditions as input and the optimal joint scheduling scheme corresponding to the rainfall conditions as the output target. Other algorithms and their combinations can also be selected according to actual needs, and no restrictions are imposed here. The specific steps in this embodiment include:
[0108] Step S9.1: Construct a Python program get_rainfall_parameter.py to calculate the characteristic parameters of the rainfall time series based on the rainfall conditions, including the cumulative rainfall duration, cumulative rainfall, cumulative rainfall before the peak, maximum 1-hour rainfall, maximum 2-hour rainfall, and maximum 4-hour rainfall;
[0109] Step S9.2: Statistically calculate the benefit evaluation indicators proposed in step S7, and retain the optimal scheduling scheme under each rainfall condition;
[0110] Step S9.3: Analyze the correlation between the rainfall characteristic parameters in Step S9.1 and the joint scheduling scheme using the Pearson correlation coefficient method, and select parameters with a correlation coefficient greater than 0.4 as training parameters for the machine learning model. The Pearson correlation coefficient can be calculated by the following formula:
[0111]
[0112] In the formula, ρ xy Let be the Pearson correlation coefficient between parameters x and y; E(x), E(y), and E(xy) are the mathematical expectations of parameters x, y, and xy, respectively; σ x , σ yLet x be the variance of parameter x and parameter y; Cov(x,y) is the covariance between parameter x and parameter y.
[0113] Step S9.4: Organize the feature parameter data selected in step S9.3, perform parameter normalization on it, and then divide the dataset into training set D and test set in a 9:1 ratio. In this embodiment, linear normalization is selected as the normalization method, which can be expressed by the following formula:
[0114]
[0115] In the formula, x' is the parameter value after normalization, x is the parameter value before normalization, min(x) is the minimum value of parameter x, and max(x) is the maximum value of parameter x.
[0116] Step S9.5: Use the training set obtained in step S9.4 to train the K-nearest neighbor model. Use the normalized parameters obtained in step S9.3 as input data, and the gate opening degree, pump station operating power, and total sewage treatment plant influent in the joint scheduling scheme as target data to train the model. The specific steps are as follows:
[0117] The training samples are represented in the format (x, f(x)); where x is the feature parameter of the sample, and x is determined by (x 1 ,x 2 ,x 3 ,…,x n Composed of ) where x n For the nth attribute value of sample x, that is, the number of feature parameters is equal to the dimension of the vector, for a new input sample x i x is calculated one by one using the Euclidean distance formula. i The distance between x and each sample in the training set is used to select samples that are similar to x. i The K nearest samples. Euclidean distance can be expressed as:
[0118]
[0119] In the formula, x i x j There are two samples respectively; Sample x i and x j The l-th eigenvalue; L(x) j ,x j ) is the sample x i and x j The distance between them.
[0120] Step S9.6: Input the test set data obtained in step S9.4 into the K-nearest neighbor model and output the optimal solution predicted by the K-nearest neighbor model;
[0121] Step S9.7: Use the training set obtained in step S9.4 to train the random forest model. Use the normalized parameters obtained in step S9.3 as input data, and the gate opening degree, pump station operating power, and total sewage treatment plant influent in the joint scheduling scheme as target data to train the model. The specific steps are as follows:
[0122] Multiple parallel training groups are generated using the Bootstrap resampling method, and the decision tree model is trained independently. First, the empirical entropy H(D) of the training set D is calculated:
[0123]
[0124] Calculate the empirical conditional entropy H(D|A) of feature A on training set D:
[0125]
[0126] Calculate information gain:
[0127] g(D,A)=H(D)-H(D|A)
[0128] Calculate the information gain ratio:
[0129]
[0130] in:
[0131]
[0132] In the formula, D is the entire training dataset, A is the feature parameter, K is the total number of categories, Ck is the k-th class, and n is the number of values for the feature parameter A.
[0133] Select the feature parameter A with the largest information gain ratio. g As a node, for feature parameter A g Possible values {a1, a2, ..., a n}, according to A in sequence g =a1,…,A=a n The training set D is split into Di and Di. i ,D2,…,D n Enter the next level, with A-{A gLet} be the set of feature parameters. Repeat the above steps until all feature parameters have been traversed, and then output the decision tree model. Further, combine all the independently generated decision tree models to construct a random forest model. When a new sample needs to be predicted, the results of each decision tree will be considered, and the final joint scheduling optimization scheme will be output through voting.
[0134] Step S9.8: Use the training set obtained in step S9.4 to train the extreme random tree model. Use the normalized parameters obtained in step S9.3 as input data, and the gate opening degree, pump station operating power, and total sewage treatment plant influent in the joint scheduling scheme as target data to train the model. The specific steps are as follows:
[0135] Select all training sets D for model training. Randomly select N feature parameters from feature parameters A. Randomly select one feature parameter as a split node. Take the split threshold with the smallest Gini coefficient as the optimal split threshold. Generate two child nodes from the data D_left and D_right. Iterate through the remaining parameters in turn.
[0136] The Gini coefficient of the training set D can be expressed as:
[0137]
[0138] The Gini coefficient of set D under characteristic parameter A can be expressed as:
[0139]
[0140] In the formula, D represents the entire training dataset, A represents the feature parameters, K represents the total number of categories, and C represents the total number of categories. k For class k, D1 and D2 are two subsets divided based on feature A.
[0141] Step S10: Verify the effect of the built model. Verify the effect of the model built in step S9 using test set data. If the accuracy requirements are met, store the model; otherwise, return to step 9 and retrain the model.
[0142] Step S11, perform joint scheduling prediction of plant, network and river: based on real-time rainfall and meteorological data, drive the model stored in step S10, and take the scheduling scheme with the smallest objective function among the three model output schemes as the final optimized scheduling scheme.
Claims
1. A method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models, characterized in that, include: Step 1: Establish a SWMM model of the target area, coupling wastewater treatment plants, urban pipe networks, and urban river and lake systems; Step 2: Calibrate the SWMM model parameters based on previous monitoring data of the target area; Step 3: Pre-design multiple plant-network-river joint scheduling schemes; Step 4: Construct an urban stormwater scenario database: Based on the rainstorm parameters of the target area and the Chicago rainfall pattern, construct short-duration extreme heavy rainfall events with different return periods to build an urban stormwater scenario database; Step 5: Construct an evaluation system for the joint scheduling effect of plants, networks, and rivers; Step 6: Construct an urban plant-network-river joint scheduling scenario database: Combine the multiple plant-network-river joint scheduling schemes pre-designed in Step 3 with the short-duration extreme heavy rainfall events constructed in Step 4. The events are combined and simulated based on the calibrated SWMM model to obtain joint dispatch scenario data for various short-duration extreme heavy rainfall events under different plant-network-river joint dispatch schemes. The combined schemes are evaluated using the plant-network-river joint dispatch effect evaluation system constructed in step 5 to form an urban plant-network-river joint dispatch scenario database. Step 7: Construct a rainfall condition-joint dispatch scheme-benefit index learning model. This model uses rainfall conditions and benefit indicators as input conditions and the optimal joint dispatch scheme as the output objective. The model is trained based on the urban plant-network-river joint dispatch scenario database constructed in step 6. Step 8: Use the trained rainfall condition-joint dispatch... The scheme-benefit index learning model obtains the optimal joint scheduling scheme for power plants, networks, and rivers under the current rainfall conditions. Step 7, constructing the rainfall condition-joint scheduling scheme-benefit index learning model, involves the following steps: Step 71, calculating rainfall time series characteristic parameters based on rainfall conditions; Step 72, obtaining the optimal scheduling scheme under each rainfall condition from the urban power plant, network, and river joint scheduling scenario database; Step 73, calculating the correlation between rainfall time series characteristic parameters and each optimal scheduling scheme, and selecting parameters with correlation coefficients greater than a preset value as training parameters for the machine learning model; Step 74, dividing the training parameters into a training set D and a test set; Step 75, establishing the machine learning model. The machine learning model is trained using a training set. The training set data serves as input, and the gate openings, pump station operating power, and total wastewater inflow at the joint scheduling scheme are used as target data. Step 75 includes a K-nearest neighbor model, a random forest model, and a limit random tree model. Step 8 selects the scheduling scheme with the smallest objective function from the three model output schemes as the final optimized scheduling scheme. The steps for training the random forest model using the training set are as follows: Multiple parallel training groups are generated using the Bootstrap resampling method, and decision tree model training is performed independently. First, the empirical entropy H(D) of the training set D is calculated. Calculate the empirical conditional entropy of feature A on training set D. : Calculate information gain: Calculate the information gain ratio: ,in: In the formula, D is the entire training dataset, A is the feature parameter, K is the total number of categories, and C is the total number of categories. k For the k-th class, n is the number of possible values for feature parameter A; select the feature parameter A with the largest information gain ratio. g As a node, for feature parameter A g The values of are {a1, a2, ..., a n }, according to A in sequence g =a1,…,A g =a n The training set D is divided into D1, D2, ..., D... n Enter the next level, with A-{A g } represents the set of feature parameters. Repeat the above steps until all feature parameters have been traversed, and then output the decision tree model. Combine all the independently generated decision tree models to construct a random forest model.
2. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 1, characterized in that, Step 1 specifically includes: Step 11, based on the river and lake system information of the target area, simplify the river into an open channel and the lake into a regulating reservoir; Step 12, connect the outlet of the pipeline to the river and lake, and set the outflow conditions according to the water level of the river and lake.
3. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 2, characterized in that, The outflow condition in step 12 is as follows: when the river water level is lower than the outlet node elevation, the outlet discharge mode is free outflow, and the flow rate can be expressed as: In the formula, A is the cross-sectional area of the discharge outlet; H0 is the water head height; The lateral contraction coefficient; is the local head loss coefficient; g is the acceleration due to gravity; when the river level is higher than the outlet, but the pipeline head is higher than the river head, the outlet discharges as a submerged outflow, and its flow rate can be expressed as: In the formula, A is the cross-sectional area of the outlet; z is the difference between the water head in the pipeline and the water head in the river. The lateral contraction coefficient; is the local head loss coefficient; g is the gravitational acceleration; when the river water level is higher than the outlet and the pipeline head is lower than the river head, the water in the pipeline will not be discharged from the outlet, and at the same time, the river water will not flow back into the pipe network.
4. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 1, characterized in that, The short-duration extreme heavy rainfall events in step 4 are specifically 1-hour, 2-hour, and 3-hour short-duration extreme heavy rainfall events.
5. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 1, characterized in that, In step 5, the evaluation indicators consist of the cumulative overflow of nodes, the cumulative number of overflowing nodes, the cumulative amount of pollutant overflow and discharge, and the cumulative amount of water pumped by the pumping station. The above indicators are comprehensively considered by constructing an objective function. The constraints are set as follows: the maximum power of the pumping station shall not exceed the rated power of the pumping station, the maximum influent flow of the sewage treatment plant shall not exceed its rated flow, the water level of the regulating lakes in the watershed shall not exceed the highest operating water level, and the water level of the flood control and drainage river shall not exceed the maximum safe water level.
6. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 1, characterized in that, The steps for training a K-nearest neighbor model using the training set are as follows: Represent the training samples as... The format in which For the feature parameters of the sample, Depend on Composition, in which For the sample The Each attribute value; for a new input sample Calculate one by one using the Euclidean distance formula The distance between the training set and each sample, and then the samples with the same distance as the training set are selected. The K nearest samples; Euclidean distance is expressed as: In the formula, , There are two samples respectively; , Samples and The l-th eigenvalue; For the sample and The distance between them.
7. The method for joint scheduling of power plants, networks, and rivers based on machine learning algorithms and rainfall-flood numerical models according to claim 1, characterized in that, The steps for training a limit random tree model using a training set are as follows: Select all training sets D for model training; randomly select N feature parameters from feature parameters A; randomly select one feature parameter as a splitting node; take the splitting threshold with the smallest Gini coefficient as the optimal splitting threshold; generate two child nodes from the data D_left and D_right; and iterate through the remaining parameters in turn. The Gini coefficient of the training set D can be expressed as: Given characteristic parameter A, the Gini coefficient of set D can be expressed as: In the formula, D is the entire training dataset, A is the feature parameter, K is the total number of categories, and C is the total number of categories. k For class k, D1 and D2 are two subsets divided based on feature A.