Reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning

Through deep learning adaptive dynamic network and reinforcement learning methods, the reservoir scheduling model is dynamically adjusted, which solves the problem of unstable prediction accuracy in traditional methods, realizes the coordinated management of reservoir flood control safety, power generation benefits and ecological protection, and improves the real-time and stability of scheduling decisions.

CN120338210BActive Publication Date: 2025-10-21水利部珠江水利委员会珠江水利综合技术中心
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510819807.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-21
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional reservoir scheduling methods are difficult to adapt to the complex and changeable hydrological and meteorological environment, resulting in unstable prediction accuracy and insufficient generalization ability. They are prone to decision-making lags and water level control errors, increasing flood control risks, reducing power generation efficiency, and destroying ecosystem stability.

Method used

A method based on deep learning adaptive dynamic network and reinforcement learning is adopted. Through real-time data fusion, adaptive dynamic Transformer network model and reinforcement learning state space, the network structure and target weights are dynamically adjusted to optimize the combination of flood discharge, power generation flow and ecological flow.

Benefits of technology

It has improved the coordinated management efficiency of reservoir flood control safety, power generation benefits and ecological protection, enhanced the prediction accuracy stability and model generalization ability, and can quickly respond to extreme weather events and reduce decision-making delays and risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338210B_ABST
    Figure CN120338210B_ABST
Patent Text Reader

Abstract

The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning relates to the technical field of reservoir scheduling and comprises the following steps: obtaining real-time water level data, meteorological data, radar image data of regional cloud and surface features of a reservoir, and performing data fusion on the three types of data; inputting the fused data into a preset adaptive dynamic Transformer network model to obtain a prediction output, wherein the prediction output is a reservoir water level sequence, a reservoir inflow sequence and a reservoir outflow sequence; constructing a reinforcement learning state space; inputting the reinforcement learning state space into a preset reinforcement learning network, wherein an action space of the reinforcement learning network comprises a flood discharge, a power generation flow and an ecological flow, and an optimization target of the reinforcement learning network is to maximize a cumulative discount reward, and a reward function is a weighted function of a flood control reward item, a power generation reward item and an ecological reward item. The method improves the collaborative management efficiency of multiple targets such as reservoir flood control safety, power generation benefit and ecological protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reservoir scheduling, and in particular to a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning. Background Art

[0002] As global climate change intensifies and extreme weather events become more frequent, traditional reservoir operation methods struggle to adapt to the complex and ever-changing hydrological and meteorological environment. Conventional operation methods are often based on fixed rules or offline historical experience and lack effective utilization of real-time meteorological and water level data. In the face of sudden rainstorms or rapid transitions between drought and flood, decision lags and water level control errors are prone to occur, thereby increasing flood control risks, reducing power generation efficiency, and undermining the stability of downstream ecosystems. Current reservoir operation methods using deep learning network structures are typically static and fixed, lacking real-time adaptive capabilities. They are unable to adjust the network structure in a timely manner as hydrological forecasting needs change dynamically, resulting in unstable prediction accuracy and insufficient model generalization.

[0003] Therefore, a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning was developed to solve the above problems. Summary of the Invention

[0004] The present invention proposes a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning to solve the problems of unstable prediction accuracy and insufficient generalization ability of existing reservoir scheduling methods.

[0005] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0006] The present invention provides a reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning, comprising:

[0007] Obtain real-time reservoir water level data, meteorological data, and radar image data of regional cloud cover and surface features and fuse the three;

[0008] The fused data is input into a preset adaptive dynamic Transformer network model to obtain a prediction output, which includes a reservoir water level sequence, an inflow sequence, and a discharge sequence;

[0009] Construct a reinforcement learning state space, which includes the predicted reservoir water level sequence, inflow sequence, outflow sequence, and the latest real-time meteorological data sequence at the future preset time step of the Transformer network model;

[0010] The reinforcement learning state space is input into a preset reinforcement learning network. The action space of the reinforcement learning network includes flood discharge, power generation flow and ecological flow. The optimization goal of the reinforcement learning network is to maximize the cumulative discounted reward. The reward function is a weighted function of flood control reward items, power generation reward items and ecological reward items. The reinforcement learning network is optimized through continuous iterative optimization and eventually converges to output the optimal combination of flood discharge, power generation flow and ecological flow.

[0011] Furthermore, the real-time water level data and meteorological data are preprocessed before data fusion. The preprocessing steps include:

[0012] Furthermore, a Kriging interpolation algorithm based on spatial distance weights was used to perform spatial interpolation processing on real-time water level data, meteorological data, and radar image data of regional cloud cover and surface features to obtain continuous data covering the entire reservoir area;

[0013] Perform outlier detection on continuous data and remove outliers in real time;

[0014] The data after removing outliers is smoothed and denoised.

[0015] Furthermore, when the error between the predicted output of the adaptive dynamic Transformer network model and the actual observation result each time exceeds a preset upper threshold, a preset number of coding layers are automatically added to the adaptive dynamic Transformer network model; when the result error is lower than a preset lower threshold, the preset number of coding layers are automatically reduced for the adaptive dynamic Transformer network model; when the error between the predicted output of the adaptive dynamic Transformer network model and the actual observation result each time is between the lower threshold and the upper threshold, the adaptive dynamic Transformer network model remains unchanged.

[0016] Furthermore, the result error is the weighted sum of the mean square error of the reservoir water level between the predicted value and the actual value, the mean square error of the inflow between the predicted value and the actual value, and the mean square error of the outflow between the predicted value and the actual value.

[0017] Furthermore, the adaptive dynamic Transformer network model outputs the weights of flood control reward items, power generation reward items and ecological reward items in the reward function based on the dynamic attention mechanism each time. .

[0018] Furthermore, according to the weights of flood control incentives, power generation incentives and ecological incentives Update target weight , the update method is:

[0019] ;

[0020] ;

[0021] ;

[0022] The above dynamic attention weights are adjusted in real time and used as the target weighting coefficient in the next reinforcement learning decision, ensuring that the system can accurately adapt to the real-time needs of different scheduling goals based on the real-time status;

[0023] The dynamic attention mechanism uses a target weight vector with trainable parameters and calculates the attention weight of each target in real time using the Softmax function.

[0024] Furthermore, R t is the cumulative discount reward, α 防洪 is the target weight coefficient of the flood control reward item, α 发电 is the target weight coefficient of the power generation reward item, α 生态 is the target weight coefficient of the ecological reward item, To predict water levels, To prevent floods, water levels are limited. is the economic benefit coefficient of power generation per unit flow, Q 发电,t is the power generation flow, Q 生态,t is the current ecological flow, Q 生态,理想 is the ideal ecological flow, R 防洪 For flood control reward items, R 发电 is the power generation reward item, R 生态 is the ecological reward item, and t is the time step.

[0025] Furthermore, the opening of each flood discharge gate of the reservoir and the output plan of the generator set are determined based on the optimal combination of the output flood discharge volume, power generation flow and ecological flow, and specific execution instructions are formed, including:

[0026] The flood discharge gate opening is calculated in real time based on the output flood discharge volume and the gate flow-opening relationship formula;

[0027] Determine the number of units started and stopped and the load distribution of each unit in real time based on the output power flow;

[0028] Construct execution instructions based on the floodgate opening, the number of units started and stopped, and the load distribution of each unit;

[0029] Determine whether the total discharge of "flood discharge + power generation" is not lower than the ecological flow. If so, it is considered that the ecological water demand has been included and no additional scheduling is required. If not, water is replenished through the ecological special sluice or low-load unit according to the difference. The difference is the difference between the ecological flow and the total discharge of "flood discharge + power generation".

[0030] Furthermore, a long-term running root mean square error historical database is constructed, and statistical analysis of the historical trends of result errors and decision errors is performed regularly. Based on the historical error trends, the preset error threshold is dynamically adjusted.

[0031] Furthermore, real-time reservoir water level data, meteorological data, and radar image data of regional clouds and surface features are obtained, including:

[0032] Acquiring real-time water level data monitored by wave-type water level sensors deployed in key sections of the reservoir area and main inflow rivers;

[0033] Obtain meteorological data from the meteorological data acquisition system, which is deployed in the reservoir area and includes precipitation sensors, wind speed sensors, and temperature and humidity sensors;

[0034] Obtain radar image data of regional cloud and surface features. Use a high-resolution radar remote sensing satellite with a spatial resolution of 30m to conduct satellite remote sensing observations of the reservoir and the upstream area of ​​the basin to obtain radar image data of regional cloud and surface features.

[0035] The beneficial effects of the present invention are:

[0036] The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning proposed in this invention solves the problem of comprehensively improving the collaborative management efficiency of multiple goals such as reservoir flood control safety, power generation efficiency and ecological protection compared with the existing technology, improving the prediction accuracy and stability, and enhancing the model generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is an overall flow chart of the reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning in the embodiment.

[0038] Figure 2 Schematic diagram of the adaptive dynamic Transformer network structure in step S2 of the embodiment.

[0039] Figure 3 This is a flowchart of the reinforcement learning multi-objective collaborative optimization module in step S3 of the embodiment.

[0040] Figure 4 This is a functional framework diagram of the intelligent collaborative scheduling platform in step S6 of the embodiment. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.

[0042] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0043] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0044] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.

[0045] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0046] like Figure 1 As shown in FIG, the flow chart of the reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning of the present invention is shown in FIG. The specific implementation method is as follows:

[0047] Step S1: Real-time data fusion and preprocessing.

[0048] In the actual implementation of this invention, a total of eight ultrasonic water level sensors (model: MB7389) were deployed within the reservoir area and at key sections of the main inflow channel. With a water level measurement accuracy of ±2cm, they collect real-time water level data every minute and upload the data to a data fusion center via the LoRa wireless transmission protocol. A meteorological data collection system consisting of a weather station was also installed within the reservoir area. Specifically, it includes a precipitation sensor (model: TB4, accuracy of ±0.5mm), a wind speed sensor (model: WindSonic, accuracy of ±0.2m / s), and a temperature and humidity sensor (model: HMP155, temperature accuracy of ±0.5°C, humidity accuracy of ±3%), collecting meteorological data every five minutes. Furthermore, a high-resolution radar remote sensing satellite with a spatial resolution of 30m is used to conduct hourly satellite remote sensing observations of the reservoir and upstream areas of the basin, acquiring radar image data of regional cloud cover and surface features in GeoTIFF format. The data fusion module first uses a Kriging interpolation algorithm based on spatial distance weights to spatially interpolate data from multiple water level measurement points and meteorological sensors, obtaining continuous data covering the entire reservoir area. Outlier detection after data fusion is implemented using the Z-Score method. Regional cloud cover and meteorological data are combined to predict sudden rainstorms in advance. Surface features are used to distinguish different runoff effects under the same rainfall amount, enabling more precise scheduling.

[0049] The specific formula is as follows:

[0050]

[0051] in, Indicates real-time measurement values, Indicates the average of historical data in the past 24 hours. represents the standard deviation, and Z represents the standard score. When the system automatically marks the data as outliers and removes them in real time. Finally, after smoothing and denoising with a Savitzky-Golay smoothing filter, the data is normalized into a sequence format and periodically updated in 10-minute units to be input to the subsequent network model.

[0052] Step S2: Accurate prediction based on adaptive dynamic network of Transformer structure.

[0053] This step uses the Transformer network structure and uses the standardized data sequence processed in step 1 as the network input to achieve short-term (next 30 minutes) accurate prediction of reservoir water level, inflow, and outflow. The specific implementation is as follows:

[0054] Step 2.1: If Figure 2As shown in the figure, the initial structure of the Transformer network is constructed: The initial structure of the Transformer network consists of 4 standard Transformer encoder layers. Each encoder layer includes a multi-head self-attention sublayer and a feedforward neural network sublayer, as well as a Layer Norm layer and residual connections. Each self-attention sublayer contains 8 attention heads, each with a dimension of 64. Therefore, the output dimension of each self-attention sublayer is 512 dimensions. The calculation formula is as follows:

[0055] The calculation method of a single attention head is:

[0056]

[0057] in, Represents the query matrix, key matrix and value matrix respectively, and the dimensions are The output of multi-head attention is:

[0058]

[0059] in, , Represents the linear transformation parameter matrix.

[0060] The feedforward neural network consists of two linear transformation layers with a structure of 512 dimensions → 2048 dimensions → 512 dimensions. The formula is as follows:

[0061]

[0062] The network's initial input data dimension is standardized data from the past 120 minutes (a total of 12 time steps, each containing characteristic data such as water level, precipitation, wind speed, temperature and humidity, for a total input dimension of 128). The output is a prediction of water level, inflow, and outflow for the next three time steps (30 minutes) (output dimension is 3D), and the weights of the flood control reward term, power generation reward term, and ecological reward term in the reward function are also output (output dimension is 3D). The Adam optimizer is used for training, with an initial learning rate of 0.001 and a loss function using the mean squared error (MSE):

[0063]

[0064] in, represents the actual observed value, represents the network output prediction value, is the number of samples in a batch. During training, the branch that outputs the weights of flood control incentives, power generation incentives, and ecological incentives is temporarily excluded from training, meaning its gradient is truncated.

[0065] Step 2.2: Adaptive dynamic adjustment of network structure: The network dynamic adjustment module is based on the error of the result calculated in real time. Automatically adjust the number of Transformer network layers, where the result error is calculated as follows:

[0066]

[0067] in, Refers to the mean square error between the predicted value and the true value output by the network, weight 、 、 Determined by offline cross-validation, Indicates the water level, Indicates the inflow flow, represents the discharge flow, t represents time, and t represents time;

[0068] Taking the latest W prediction cycles as the sliding window, first calculate the comprehensive error sequence The sliding mean of and standard deviation , the standard deviation takes the form of a weighted covariance:

[0069]

[0070] in The indicator weight vector determined for offline cross-validation, is the covariance matrix of the three types of RMSE within the window, and then the threshold is dynamically constructed:

[0071] ;

[0072] ;

[0073] The relaxation coefficient Through Bayesian optimization, it is determined offline on the historical data set. When the overall result error gradually decreases, Then it shrinks, On the contrary, when the result error increases sharply, It is dynamically lifted to avoid misjudgment, and once the real-time result error Beyond the new upper threshold, the layer-increasing operation will still be triggered. When the error between the predicted result and the actual observation result exceeds the upper threshold, When the error of the result is lower than the lower threshold, the number of network layers increases by 1, and the maximum increase does not exceed 3; When , the number of network layers is reduced by 1, and the specific adjustment rules are as follows:

[0074] ;

[0075] The dynamic adjustment mechanism of the number of network layers can effectively balance the computing load and prediction accuracy, maintaining the real-time and stability of the prediction.

[0076] Step 2.3: Implementation of dynamic attention mechanism: In this step, a dynamic attention mechanism is designed to adaptively adjust the prediction weights of three different goals: flood control, power generation, and ecology. Update target weight , the update method is:

[0077] ;

[0078] ;

[0079] ;

[0080] The above dynamic attention weights are adjusted in real time and used as the target weighting coefficient in the next reinforcement learning decision, ensuring that the system can accurately adapt to the real-time needs of different scheduling goals based on the real-time status;

[0081] The dynamic attention mechanism uses a target weight vector with trainable parameters and calculates the attention weight of each target in real time using the Softmax function.

[0082] After the above dynamic attention weights are adjusted in real time, they serve as the target weighting coefficient in the next reinforcement learning decision, ensuring that the system can accurately adapt to the real-time needs of different scheduling goals according to the real-time status.

[0083] The dynamic attention mechanism uses a target weight vector with trainable parameters and calculates the attention weight of each target in real time using the Softmax function.

[0084] Step 3: Implementation of the reinforcement learning multi-objective collaborative optimization module.

[0085] Step 3.1: Define the reinforcement learning state space for multi-objective optimization, which includes meteorological data, the predicted water level sequence output by the adaptive network, the inflow sequence, and the outflow sequence. Figure 3 As shown, the state space of reinforcement learning is constructed. Based on the prediction results output by the dynamic Transformer network in step 2, the state vector of the reinforcement learning module is defined as S t Specifically, the state vector contains the predicted water level sequence for the next three time steps output by the Transformer network. , Inbound flow forecast sequence , and the downstream flow prediction sequence Q out,t =[Q out,t+1 ,Q out,t+2,Q out,t+3 ] and the latest real-time meteorological data series M t , meteorological data includes precipitation, wind speed, temperature and humidity. After the above state data are spliced, the state space dimension is controlled between 100-500 dimensions, and after standardization, it is input into the reinforcement learning network.

[0086] Step 3.2: Define the reinforcement learning action space. Action vector a t It includes three variables: flood discharge, power generation flow, and ecological flow, and the specific settings are as follows:

[0087] Flood discharge adjustment range: 50 to 5000 m³ / s, step size 50 m³ / s;

[0088] Power generation flow adjustment range: 100 to 2000 m³ / s, step size 20 m³ / s;

[0089] Ecological flow regulation range: 10 to 500 m³ / s, step size 10 m³ / s;

[0090] Discrete combinations of the action space are generated using a grid search method, and the state vector is mapped to the action space through a three-layer fully connected network, so that the reinforcement learning model can perform efficient action selection.

[0091] Step 3.3: Specific implementation of the reinforcement learning network. The reinforcement learning network uses a deep Q network, and the network structure is as follows:

[0092] Input layer: input state vector S t ;

[0093] Hidden layer: Consists of 4 convolutional layers and 2 fully connected layers. The convolutional kernel sizes are [8×8, 4×4, 3×3, 3×3], the convolution strides are [4, 2, 1, 1], and the output feature dimensions are 128, 128, 64, and 64, respectively. It then passes through two fully connected layers. The first layer outputs a 128-dimensional feature vector using the ReLU activation function, and the second layer outputs the Q value of the action space.

[0094] Output layer: Q value of each action combination.

[0095] The optimization goal of the reinforcement learning network is to maximize the cumulative discounted reward , the reward function is defined as the weighted result of the three objectives (flood control, power generation, and ecology). The specific formula is:

[0096] ;

[0097] Among them, the reward weight α 防洪 , α 发电 , α生态 The calculation results of the dynamic attention mechanism from step 2.3 are updated dynamically in real time. Specifically, the target rewards are defined as follows:

[0098] Flood Control Rewards 防洪 Negative feedback is given according to the degree to which the predicted water level exceeds the flood control limit water level, which is defined as: .in, To predict water levels, To limit the water level for flood control.

[0099] Power Generation Incentive Item R 发电 The economic benefits calculated based on the current power generation flow are defined as: Among them, k 发电 is the economic benefit coefficient of power generation per unit flow, Q 发电,t is the power generation flow;

[0100] Ecological Reward Item R 生态 Calculation is based on the degree to which the downstream ecological flow deviates from the ideal ecological flow: Where Q 生态,t is the current ecological flow, Q 生态,理想 It is the ideal ecological flow.

[0101] During training, the parameters of the dynamic Transformer network are frozen, only the layers with output weights are activated, and the training is performed jointly with the reinforcement learning network.

[0102] Through continuous iterative optimization, the reinforcement learning network eventually converges to the optimal strategy for multi-objective collaborative optimization and outputs the optimal combination solution for each flow.

[0103] The network update learning rate of the reinforcement learning network is between 0.001 and 0.005; the dynamic range of the weight coefficient of each objective in the reward function is 0.4-0.8 for flood control, 0.1-0.5 for power generation, and 0.1-0.3 for ecology, and the target weight is updated every 30 minutes.

[0104] Step 4: Implementation of real-time reservoir operation decision module.

[0105] In this embodiment, the real-time reservoir scheduling decision module uses the optimal strategy output by the reinforcement learning multi-objective collaborative optimization module to determine the opening of each flood discharge gate of the reservoir and the output plan of the generator set in real time, and form specific execution instructions. Specifically, the flood discharge gate opening control instruction Flood discharge output by the reinforcement learning module Obtained through real-time calculation of the gate flow-opening relationship formula:

[0106]

[0107] in, It is the empirical relationship function between flood discharge and opening obtained based on the actual measurement of the hydraulic characteristics of the gate.

[0108] The output adjustment command of the generator set is based on the power generation flow output by reinforcement learning Determine the number of units started and stopped and the load distribution of each unit in real time:

[0109]

[0110] in, is the power generation capacity of the i-th unit, is the water density, is the acceleration due to gravity, H is the water head, is the real-time water head, is the unit efficiency, is the number of starting units.

[0111] On this basis, check "Q 泄洪,t +Q 发电,t "Is the total discharge volume no less than Q 生态,t If it is satisfied, it is considered that the ecological water demand has been included and no additional scheduling is required; if it is insufficient, the difference Q 补生态,t =Q 生态,t- (Q 泄洪,t +Q 发电,t ) Water is added through ecological special sluice holes or low-load units to ensure that the downstream ecological flow meets the standards, so as to supplement the ecological flow.

[0112] Under extreme rainfall conditions (such as rainfall intensity in the reservoir area exceeding 30 mm / h), the system automatically increases the flood control target weight α 防洪 Improved to 0.85, the reinforcement learning network outputs high-intensity flood discharge strategies in real time, quickly responding to rainstorm events to ensure that the water level in the reservoir area is always strictly controlled below the flood control limit to avoid flood risks.

[0113] The dispatch decision response delay is controlled within 1 minute; under extreme rainfall conditions, the system automatically increases the flood control target weight to between 0.7-0.9, and adjusts the flood discharge volume as quickly as possible to ensure that the reservoir water level is always controlled below the flood control limit.

[0114] Step 5: Implement rolling optimization and feedback update mechanism.

[0115] Step 5.1: Evaluate the error of the results in real time Updated with dynamic network. The system automatically calculates the result error once every hour , and calculate the upper and lower thresholds respectively according to step S2.2, and trigger the update process of Transformer adaptive dynamic network structure and parameters according to the adjustment rules to reduce the subsequent result error The network update is achieved through the gradient descent algorithm with a learning rate of 0.001. After the parameter adjustment is completed, it is immediately put back into the real-time prediction work.

[0116] Step 5.2: Regularly fine-tune the reinforcement learning strategy. At the end of each day, the system automatically compiles statistics on the multi-objective optimization results of the reinforcement learning decision-making, including flood risk reduction, power generation revenue, and ecological flow protection, and compares and analyzes them with the goals of the previous cycle. At the end of each week, fine-tune the parameters of the reinforcement learning network based on the aggregated data. This fine-tuning process utilizes an experience replay strategy, storing the state-action-reward data from the past week. The parameters of the reinforcement learning network are then updated through batch random sampling to ensure continuous network optimization and generalization performance.

[0117] Step 5.3: Establishment and regular optimization of long-term error database. The system establishes a long-term error history database and regularly optimizes the result errors every month. , and conduct statistical analysis on the historical trend of decision errors. Based on the historical error trend, dynamically adjust the prediction error threshold (initial 0.05m, adjust the range of ±0.01m according to the trend) and the reinforcement learning target weight 、 、 The dynamic adjustment range is controlled within ±10% to ensure stability and adaptability during long-term operation.

[0118] Step 6: Deploy and implement the intelligent collaborative scheduling platform.

[0119] like Figure 4 As shown, the intelligent collaborative scheduling platform is implemented through an integrated intelligent management platform deployed on high-performance servers. This platform features functional modules such as real-time data monitoring, automated prediction and decision-making, real-time alarms, and historical data analysis. The platform utilizes a distributed architecture and is deployed on high-performance Intel Xeon servers (32-core CPUs, 128GB of memory). It features disaster recovery and remote web access. The real-time data monitoring module receives and displays sensor data such as water level and weather conditions. The adaptive network prediction module updates predictions every 10 minutes. The reinforcement learning decision-making module outputs optimized scheduling instructions in real time and automatically transmits them to on-site execution equipment. The automatic scheduling execution module executes gate opening and power generation load in real time. The system's alarm mechanism responds within 30 seconds and automatically detects events such as excessive water levels and extreme weather anomalies, triggering real-time alarms that are sent to management personnel via SMS and email. The platform's visual interface includes real-time reservoir status (such as real-time water level curves and gate opening indicators), weather trend charts, and multi-objective optimization decision charts. It also supports rapid query and analysis of historical data, providing comprehensive technical support for integrated reservoir management through charts and reports.

[0120] Step 7: Implement long-term system operation status monitoring and performance evaluation.

[0121] In order to ensure the long-term stability and adaptability of the system, a comprehensive long-term operation status monitoring and performance evaluation mechanism has been established to regularly evaluate the overall performance of the system. A comprehensive system operation evaluation report is generated every quarter, and the evaluation indicators include:

[0122] Flood risk reduction rate:

[0123] ;

[0124] Power generation efficiency improvement rate:

[0125] ;

[0126] Ecological flow guarantee rate:

[0127] ;

[0128] The evaluation error for each metric is controlled within ±5%. Based on the quarterly evaluation results, if any metric shows a significant downward trend (a drop exceeding 10%), the adaptive dynamic network structure and reinforcement learning network parameters are automatically updated and adjusted to restore system performance. Furthermore, a system operation log and abnormal event recording mechanism has been established to provide long-term storage and management of scheduling strategy changes, predicted abnormalities, and parameter adjustment records, forming a comprehensive operational database for annual technology upgrade and system maintenance decisions.

[0129] The present invention comprehensively improves the efficiency of collaborative management of multiple objectives such as reservoir flood control safety, power generation efficiency and ecological protection through real-time data fusion preprocessing technology, adaptive dynamic network structure design and multi-objective reinforcement learning decision-making mechanism. Especially under extreme weather conditions, the present invention can quickly respond and dynamically adjust the scheduling target weights, effectively reducing the decision-making delays and risks in traditional reservoir scheduling methods. Through the rolling update and feedback optimization mechanism of dynamic network structure and reinforcement learning decision-making, the present invention can maintain high adaptability and stability for a long time, meeting the strict requirements of actual reservoir management for real-time and accuracy. This method can be widely used in the intelligent and refined management of large and medium-sized reservoirs, significantly improving the level of reservoir scheduling decision-making and the ability to ensure safe operation.

[0130] The present invention provides a reservoir scheduling method based on deep learning adaptive dynamic networks and reinforcement learning. It performs real-time fusion and precise preprocessing of multi-source data through real-time water level sensors, meteorological data acquisition systems, and high-resolution satellite remote sensing equipment, effectively improving the accuracy and integrity of the real-time status data of the reservoir. By designing an adaptive dynamic network and dynamic attention mechanism based on the Transformer structure, the present invention achieves short-term and precise prediction of reservoir water level, inflow, and outflow, and can automatically and dynamically adjust the network structure according to the result error to ensure a good balance between prediction accuracy and computational efficiency. At the same time, the present invention combines deep reinforcement learning methods to construct a multi-objective collaborative optimization decision-making module, which outputs the optimal combination strategy of flood discharge, power generation, and ecological flow in real time based on dynamic prediction results, thereby achieving real-time dynamic intelligent collaborative regulation of multiple objectives of the reservoir. In the real-time scheduling execution link, the present invention can quickly convert the optimization strategy into precise execution instructions for gate opening and generator output, respond quickly under extreme weather conditions, and ensure flood control safety.

[0131] The present invention proposes a complete and innovative rolling optimization feedback update mechanism, which automatically triggers network parameter fine-tuning and strategy update through real-time evaluation of result errors and reinforcement learning decision effects, effectively improving the stability and adaptability of long-term operation. The deployment of the intelligent collaborative scheduling platform realizes real-time data monitoring, automated execution of prediction decisions, and rapid alarm of abnormal events, effectively improving the automation, refinement and intelligence level of reservoir management. In addition, the present invention further establishes a monitoring and evaluation mechanism for the long-term operation status and performance of the system, which can regularly evaluate the realization of flood control safety, power generation benefits and ecological protection goals, and adaptively optimize the network structure and parameters based on the evaluation results, so that the long-term operation performance of the system remains in the optimal state. Experimental verification in an actual reservoir environment shows that the implementation of the present invention can significantly reduce the flood control risk of the reservoir, improve the economic benefits of power generation and effectively protect the ecological flow. It has the outstanding advantages of rapid real-time response, strong multi-objective collaborative optimization capabilities, and high long-term operation stability. It can effectively meet the management needs of modern reservoirs in complex meteorological and hydrological environments.

[0132] The reservoir scheduling method based on deep learning adaptive dynamic network proposed in the present invention solves the problem.

[0133] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning, characterized by: include: Obtain real-time reservoir water level data, meteorological data, and radar image data of regional cloud cover and surface features and fuse the three; The fused data is input into the preset adaptive dynamic Transformer network model to obtain the prediction output, which includes the reservoir water level series, the inflow flow series and the outflow flow series. Construct a reinforcement learning state space, which includes the predicted reservoir water level sequence, inflow sequence, outflow sequence, and the latest real-time meteorological data sequence at the future preset time step of the Transformer network model; The reinforcement learning state space is input into a pre-set reinforcement learning network. The action space of the reinforcement learning network includes flood discharge, power generation flow, and ecological flow. The optimization goal of the reinforcement learning network is to maximize the cumulative discounted reward. The reward function is a weighted function of the flood control reward item, the power generation reward item, and the ecological reward item. Through continuous iterative optimization, the reinforcement learning network eventually converges and outputs the optimal combination of flood discharge, power generation flow, and ecological flow. When the error between the predicted output of the adaptive dynamic Transformer network model and the actual observation result exceeds a preset upper threshold, the adaptive dynamic Transformer network model automatically adds a preset number of coding layers. When the result error is lower than a preset lower threshold, the adaptive dynamic Transformer network model automatically reduces the preset number of coding layers. When the error between the predicted output of the adaptive dynamic Transformer network model and the actual observation result is between the lower threshold and the upper threshold, the adaptive dynamic Transformer network model remains unchanged. The result error is the weighted sum of the mean square error of the reservoir water level between the predicted value and the actual value, the mean square error of the inflow between the predicted value and the actual value, and the mean square error of the outflow between the predicted value and the actual value. Adaptive dynamic Transformer network model adaptive dynamic adjustment according to the result error of real-time calculation Automatically adjust the number of Transformer network layers, where the result error is calculated as follows: , in, Refers to the mean square error between the predicted value and the true value output by the network, weight 、 、 Determined by offline cross-validation, H represents the water level, Indicates the inflow flow, represents the discharge flow, and t represents the time; Taking the latest W prediction cycles as the sliding window, first calculate the comprehensive error sequence The sliding mean of and standard deviation , the standard deviation takes the form of a weighted covariance: , in The indicator weight vector determined for offline cross-validation, is the covariance matrix of the three types of RMSE within the window, and then the threshold is dynamically constructed: ; ; The relaxation coefficient Through Bayesian optimization, it is determined offline on the historical data set. When the overall result error gradually decreases, Then it shrinks, On the contrary, when the result error increases sharply, It is dynamically lifted to avoid misjudgment, and once the real-time result error Beyond the new upper threshold, the layer-increasing operation will still be triggered. When the error between the predicted result and the actual observation result exceeds the upper threshold, When the error of the result is lower than the lower threshold, the number of network layers increases by 1, and the maximum increase does not exceed 3; When , the number of network layers is reduced by 1, and the specific adjustment rules are as follows: ; Based on the optimal combination of flood discharge, power generation flow and ecological flow, the opening of each flood discharge gate of the reservoir and the output plan of the generator set are determined, and specific execution instructions are formed, including: The flood discharge gate opening is calculated in real time based on the output flood discharge volume and the gate flow-opening relationship formula; Determine the number of units started and stopped and the load distribution of each unit in real time based on the output power flow; Construct execution instructions based on the floodgate opening, the number of units started and stopped, and the load distribution of each unit; Determine whether the total discharge of "flood discharge + power generation" is not lower than the ecological flow. If so, it is considered that the ecological water demand has been included and no additional scheduling is required. If not, water is replenished through the ecological dedicated sluice or low-load units according to the difference. The difference is the difference between the ecological flow and the total discharge of "flood discharge + power generation".

2. The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1 is characterized in that: Before data fusion, the real-time water level data and meteorological data are preprocessed. The preprocessing steps include: Using the Kriging interpolation algorithm based on spatial distance weights, real-time water level data, meteorological data, and radar image data of regional cloud cover and surface features are spatially interpolated to obtain continuous data covering the entire reservoir area. Perform outlier detection on continuous data and remove outliers in real time; The data after removing outliers is smoothed and denoised.

3. The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1 is characterized in that: The adaptive dynamic Transformer network model outputs the weights of flood control reward items, power generation reward items, and ecological reward items in the reward function based on the dynamic attention mechanism each time. .

4. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 3 is characterized in that: According to the weights of flood control incentives, power generation incentives and ecological incentives Update target weight , the update method is: ; ; ; The dynamic attention mechanism uses a target weight vector with trainable parameters and calculates the attention weight of each target in real time using the Softmax function.

5. The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning according to claim 4 is characterized in that: The formula of the reward function is as follows: ; ; ; ; R t is the cumulative discount reward, α 防洪 is the target weight coefficient of the flood control reward item, α 发电 is the target weight coefficient of the power generation reward item, α 生态 is the target weight coefficient of the ecological reward item, To predict water levels, To prevent floods, water levels are limited. is the economic benefit coefficient of power generation per unit flow, Q 发电,t is the power generation flow, Q 生态,t is the current ecological flow, Q 生态,理想 is the ideal ecological flow, R 防洪 For flood control reward items, R 发电 is the power generation reward item, R 生态 It is an ecological reward item.

6. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1 is characterized in that: Build a long-term running root mean square error historical database, regularly conduct statistical analysis on the historical trends of result error and decision error, and dynamically adjust the preset error threshold based on the historical error trends.

7. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1 is characterized in that: Obtain real-time reservoir water level data, meteorological data, and radar imagery of regional cloud cover and surface features, including: Acquiring real-time water level data monitored by wave-type water level sensors deployed in key sections of the reservoir area and main inflow rivers; Obtain meteorological data from the meteorological data acquisition system, which is deployed in the reservoir area and includes precipitation sensors, wind speed sensors, and temperature and humidity sensors; Obtain radar image data of regional cloud and surface features. Use a high-resolution radar remote sensing satellite with a spatial resolution of 30m to conduct satellite remote sensing observations of the reservoir and the upstream area of ​​the basin to obtain radar image data of regional cloud and surface features.

Citation Information

Patent Citations

  • Reservoir group joint optimization scheduling method based on MADDPG reinforcement learning

    CN115952958A

  • Flood warning method and system integrating meteorological and hydrological sensitivity

    CN117575873A

  • Method and system for predicting power generation capacity of hydropower station

    CN119721368A