Reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning

Through deep learning adaptive dynamic network and reinforced learning reservoir scheduling method, real-time data fusion and preprocessing of reservoirs are realized, and an adaptive Transformer network model is built to optimize the combination of flood discharge, power generation and ecological flow, solving the problems of unstable prediction accuracy and insufficient generalization capabilities of traditional reservoir scheduling methods, and improving the efficiency of flood control safety, power generation benefits and ecological protection of reservoirs.

CN120338210AActive Publication Date: 2025-07-18水利部珠江水利委员会珠江水利综合技术中心

Patent Information

Application Number
CN202510819807.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional reservoir scheduling methods are difficult to adapt to the complex and changeable hydrological and meteorological environment, resulting in unstable prediction accuracy and insufficient generalization capabilities, which are prone to decision-making lag and water level control errors, increasing flood control risks, reducing power generation efficiency, and undermining ecosystem stability.

Method used

The reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning is adopted. Through real-time data fusion and preprocessing, an adaptive Transformer network model is built, combined with reinforcement learning state space and dynamic attention mechanism, and the combination of flood discharge, current generation and ecological flow is optimized to achieve multi-objective collaborative optimization.

Benefits of technology

It improves the efficiency of coordinated management of reservoir flood control safety, power generation benefits and ecological protection, improves prediction accuracy and stability and model generalization capabilities, and can quickly respond to extreme weather events and reduce decision-making delays and risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338210A_ABST
    Figure CN120338210A_ABST
Patent Text Reader

Abstract

The invention discloses a reservoir scheduling method based on a deep learning adaptive dynamic network and reinforcement learning, which relates to the technical field of reservoir scheduling and comprises the following steps: acquiring real-time water level data, meteorological data and radar image data of regional cloud and earth surface characteristics of a reservoir and performing data fusion on the three data; the fused data are input into a preset self-adaptive dynamic Transform network model, prediction output is obtained, and the prediction output is a reservoir water level sequence, a reservoir inflow sequence and a discharged flow sequence; constructing a reinforcement learning state space; the reinforcement learning state space is input into a preset reinforcement learning network, an action space of the reinforcement learning network comprises flood discharge, power generation flow and ecological flow, an optimization target of the reinforcement learning network is maximized accumulated discount rewards, and a reward function is a weighting function of a flood prevention reward item, a power generation reward item and an ecological reward item. According to the invention, the cooperative management efficiency of multiple targets such as reservoir flood control safety, power generation benefits and ecological protection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reservoir scheduling, and in particular to a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning. Background Art

[0002] As global climate change intensifies and extreme weather events occur frequently, traditional reservoir scheduling methods are difficult to adapt to the complex and changing hydrological and meteorological environment. Conventional scheduling methods are often based on fixed rules or offline historical experience, lacking effective use of real-time meteorological and water level data. In the face of sudden rainstorms or rapid changes from drought to flood, decision lags and water level control errors are prone to occur, thereby increasing flood control risks, reducing power generation efficiency, and undermining the stability of downstream ecosystems. The reservoir scheduling methods currently used in deep learning network structures are usually static fixed structures, lacking real-time adaptive capabilities, and unable to adjust the network structure in time with the dynamic changes in hydrological forecasting needs, resulting in unstable prediction accuracy and insufficient model generalization capabilities.

[0003] Therefore, a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning was developed to solve the above problems. Summary of the invention

[0004] The present invention proposes a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning to solve the problems of unstable prediction accuracy and insufficient generalization ability of existing reservoir scheduling methods.

[0005] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0006] The reservoir dispatching method based on deep learning adaptive dynamic network and reinforcement learning of the present invention comprises:

[0007] Obtain the real-time water level data of the reservoir, meteorological data, and radar image data of regional clouds and surface features and fuse the three data;

[0008] The fused data is input into a preset adaptive dynamic Transformer network model to obtain a prediction output, wherein the prediction output includes a reservoir water level sequence, an inflow sequence, and a discharge sequence;

[0009] Construct a reinforcement learning state space, which includes the predicted reservoir water level sequence, inflow sequence, outflow sequence and the latest real-time meteorological data sequence at the future preset time step of the Transformer network model;

[0010] Input the reinforcement learning state space into a preset reinforcement learning network. The action space of the reinforcement learning network includes flood discharge, power generation flow rate, and ecological flow rate. The optimization objective of the reinforcement learning network is to maximize the cumulative discounted reward. The reward function is a weighted function of the flood control reward term, power generation reward term, and ecological reward term. The reinforcement learning network converges to output the optimal combination of flood discharge, power generation flow rate, and ecological flow rate through continuous iterative optimization.

[0011] Furthermore, preprocess the real-time water level data and meteorological data before data fusion. The preprocessing steps include:

[0012] Furthermore, use the Kriging interpolation algorithm based on spatial distance weights to perform spatial interpolation on the real-time water level data, meteorological data, radar image data of regional cloud cover and surface features, to obtain continuous data covering the entire reservoir area;

[0013] Detect outliers in the continuous data and remove them in real time;

[0014] Perform smoothing and denoising on the data after removing outliers.

[0015] Furthermore, when the prediction output of the adaptive dynamic Transformer network model each time exceeds the preset upper threshold compared with the actual observation result, automatically add a preset number of encoding layers to the adaptive dynamic Transformer network model. When the result error is lower than the preset lower threshold, automatically reduce a preset number of encoding layers from the adaptive dynamic Transformer network model. When the prediction output of the adaptive dynamic Transformer network model each time is between the lower threshold and the upper threshold compared with the actual observation result, the adaptive dynamic Transformer network model remains unchanged.

[0016] Furthermore, the result error is the weighted sum of the mean square error of the predicted and actual reservoir water levels, the mean square error of the predicted and actual inflow rates, and the mean square error of the predicted and actual outflow rates of the reservoir.

[0017] Furthermore, the adaptive dynamic Transformer network model outputs the weights of the flood control reward term, power generation reward term, and ecological reward term in the reward function based on the dynamic attention mechanism each time.

[0018] Furthermore, according to the weights of the flood control reward term, power generation reward term, and ecological reward term Update the target weights , and the update method is: ; ;

[0019] ;

[0020] After the above real-time adjustment of the dynamic attention weights, as the target weighting coefficient in the next step of reinforcement learning decision-making, it ensures that the system can accurately adapt to the real-time needs of different scheduling objectives according to the real-time state;

[0021] The dynamic attention mechanism adopts a target weight vector with trainable parameters to calculate the attention weights of each target in real time using the Softmax function.

[0022] Furthermore, the formula of the reward function is as follows:

[0023] ;

[0024] ;

[0025] ;

[0026] ;

[0027] R t is the cumulative discounted reward, α 防洪 is the weight coefficient of the flood control reward item, α 发电 is the weight coefficient of the power generation reward item, α 生态 is the weight coefficient of the ecological reward item, H t is the predicted water level, H 限 is the flood control limit water level, k 发电 is the economic benefit coefficient of power generation per unit flow, Q 发电,t is the power generation flow, Q 生态,t is the current ecological flow, Q 生态,理想 is the ideal ecological flow, and t is the time step.

[0028] Furthermore, according to the optimal combination plan of the flood discharge amount, power generation flow, and ecological flow output, determine the opening degree of each flood discharge gate of the reservoir and the output plan of the generator sets, and form specific execution instructions, which specifically include:

[0029] Calculate the opening degree of the flood discharge gate in real time according to the output flood discharge amount and the gate flow-opening relationship formula;

[0030] Determine the number of generator sets started and stopped and the load distribution of each generator set in real time according to the output power generation flow;

[0031] Construct execution instructions according to the opening degree of the flood discharge gate, the number of generator sets started and stopped, and the load distribution of each generator set;

[0032] Judge whether the total discharge of "flood discharge + power generation" is not lower than the ecological flow. If so, it is considered that the ecological water demand has been included and no additional scheduling is required. If not, make up the difference through the ecological special sluice holes or low-load units. The difference is the difference between the ecological flow and the total discharge of "flood discharge + power generation".

[0033] Furthermore, construct a historical database of the root mean square error of long-term operation, regularly conduct statistical analysis on the historical trends of prediction errors and decision-making errors, and dynamically adjust the preset error threshold based on the historical error trends.

[0034] Furthermore, obtain the real-time water level data of the reservoir, meteorological data, and radar image data of the regional cloud layer and surface characteristics, including:

[0035] Obtain the real-time water level data monitored by the wave-type water level sensor, and the wave-type water level sensor is deployed at key cross-sections of the reservoir area and the main incoming river channels;

[0036] Obtain the meteorological data of the meteorological data acquisition system. The meteorological data acquisition system is deployed in the reservoir area and includes a precipitation sensor, a wind speed sensor, and a temperature and humidity sensor;

[0037] Obtain the radar image data of the regional cloud layer and surface characteristics. Use a high-resolution radar remote sensing satellite with a spatial resolution of 30m to conduct satellite remote sensing observations on the reservoir and the upstream area of the basin to obtain the radar image data of the regional cloud layer and surface characteristics.

[0038] The beneficial effects of the present invention are as follows:

[0039] The reservoir scheduling method based on the deep learning adaptive dynamic network and reinforcement learning proposed by the present invention solves the problem that compared with the prior art, it comprehensively improves the collaborative management efficiency of multiple objectives such as reservoir flood control safety, power generation benefits, and ecological protection, improves the prediction accuracy stability, and enhances the model generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the overall flowchart of the reservoir scheduling method based on the deep learning adaptive dynamic network and reinforcement learning in the embodiment.

[0041] Figure 2 It is the schematic diagram of the adaptive dynamic Transformer network structure in step S2 of the embodiment.

[0042] Figure 3 It is the flowchart of the reinforcement learning multi-objective collaborative optimization module in step S3 of the embodiment.

[0043] Figure 4 It is the functional framework diagram of the intelligent collaborative scheduling platform in step S6 of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0045] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0046] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0047] In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0048] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.

[0049] As Figure 1 shown, the flowchart of the reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning of the present invention is as follows:

[0050] Step S1: Real-time data fusion and preprocessing.

[0051] In the actual implementation of the present invention, a total of 8 ultrasonic water level sensors (model: MB7389) are deployed at key sections of the reservoir area and the main incoming river channels. The water level measurement accuracy is ±2 cm. Real-time water level data is collected every 1 minute and uploaded to the data fusion center in real time through the LoRa wireless transmission protocol. At the same time, a meteorological data acquisition system composed of weather stations is installed in the reservoir area, specifically including: a precipitation sensor (model: TB4, accuracy ±0.5 mm), a wind speed sensor (model: WindSonic, accuracy ±0.2 m / s), and a temperature and humidity sensor (model: HMP155, temperature accuracy ±0.5 °C, humidity accuracy ±3%). Meteorological data is collected every 5 minutes. In addition, a high-resolution radar remote sensing satellite with a spatial resolution of 30 m is used to conduct satellite remote sensing observations on the reservoir and the upstream area of the basin every hour to obtain radar image data of the regional cloud layer and surface features in the GeoTIFF data format. The data fusion module first uses the Kriging interpolation algorithm based on spatial distance weights to perform spatial interpolation processing on the data of multiple water level measurement points and meteorological sensors to obtain continuous data covering the entire reservoir area. The outlier detection after data fusion is realized based on the Z-Score method. The regional cloud layer is combined with meteorological data to predict sudden heavy rain and other situations in advance; the surface features are used to distinguish different runoff generation effects under the same rainfall amount for more accurate scheduling.

[0052] The specific formula is as follows:

[0053]

[0054] Among them, x represents the real-time measured value, μ represents the mean value of the historical data in the past 24 hours, σ represents the standard deviation, and Z represents the standard score. When |Z| > 3, the system automatically marks the data as an outlier and deletes it in real time. Finally, after smoothing and denoising processing by the Savitzky-Golay smoothing filter, the data is standardized into a sequence form and updated to the input of the subsequent network model periodically in units of 10 minutes.

[0055] Step S2: Accurate prediction based on the adaptive dynamic network of the Transformer structure.

[0056] In this step, the Transformer network structure is adopted, and the standardized data sequence processed in step 1 is used as the input of the network to achieve short-term (next 30 minutes) accurate prediction of the reservoir water level, incoming flow, and discharge flow. The specific implementation method is as follows:

[0057] Step 2.1: As Figure 2As shown in the figure, the initial structure of the Transformer network is constructed as follows: The initial structure of the Transformer network consists of 4 standard Transformer encoder layers. Each encoder layer includes a multi-head self-attention sub-layer, a feed-forward neural network sub-layer, as well as a Layer Norm layer and a residual connection. Each self-attention sub-layer contains 8 attention heads, and the dimension of each head is 64. Therefore, the output dimension of each self-attention sub-layer is 512 dimensions, and the calculation formula is as follows:

[0058] The calculation method of a single attention head is:

[0059]

[0060] Among them, Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix, and their dimensions are all . The output of the multi-head attention is:

[0061]

[0062] Among them, h = 8, represents the linear transformation parameter matrix.

[0063] The feed-forward neural network consists of two linear transformation layers, and the structure is 512 dimensions → 2048 dimensions → 512 dimensions. The formula is as follows:

[0064]

[0065] The dimension of the initial input data of the network is the standardized data of the past 120 minutes (a total of 12 time steps, each time step contains feature data such as water level, precipitation, wind speed, temperature and humidity, etc., and the total input dimension is 128 dimensions). The output predicts the water level, inflow and outflow of the next 3 time steps (30 minutes) (the output dimension is 3 dimensions), and at the same time outputs the weights of the flood control reward item, power generation reward item and ecological reward item in the reward function (the output dimension is 3 dimensions). The Adam optimizer is used in the training process, and the initial learning rate is 0.001. The loss function uses the mean square error (MSE):

[0066]

[0067] Among them, represents the actual observed value, represents the predicted value of the network output, and N is the number of batch samples. During the training process, the branches that output the weights of the flood control reward item, power generation reward item and ecological reward item do not participate in the training, that is, their gradients are truncated.

[0068] Step 2.2: Adaptive dynamic adjustment of the network structure: The network dynamic adjustment module calculates the prediction error Automatically adjust the number of Transformer network layers. Among them, the prediction error is calculated as follows:

[0069]

[0070] Among them, RMSE refers to the mean square error between the predicted value and the true value output by the network, and the weights 、 、 are determined by offline cross-validation.

[0071] Taking the last W prediction cycles as a sliding window, first calculate the sliding mean of the comprehensive error sequence and the standard deviation . To take into account the dimensional differences and correlations of the output errors of the water level, inflow, and outflow, the standard deviation is in the form of weighted covariance:

[0072]

[0073] Among them is the index weight vector determined by offline cross-validation, is the covariance matrix of the three types of RMSE within the window. Subsequently, a threshold is dynamically constructed based on the idea of "empirical 3σ confidence band":

[0074] ;

[0075] ;

[0076] Among them, the relaxation coefficient is determined offline on the historical dataset through Bayesian optimization to balance the error fluctuation and the error convergence speed. The resulting adaptive threshold can shrink or relax in real time with seasonal hydrological changes, the model's self-convergence process, and sudden extreme events: when the overall prediction error gradually decreases, shrinks accordingly, is adjusted downward accordingly, prompting the number of network layers to automatically decrease while ensuring accuracy; conversely, when abnormal situations such as continuous heavy rain cause the error to increase sharply, is dynamically raised to avoid misjudgment, and once the real-time error exceeds the new upper threshold, the layer addition operation will still be triggered to quickly improve the model's expression ability. When the error between each prediction result and the actual observation result exceeds the upper threshold , the number of network layers is increased by 1, and the maximum increase does not exceed 3; when the error is lower than the lower threshold , the number of network layers is decreased by 1, and the specific adjustment rule is:

[0077]

[0078] The dynamic adjustment mechanism of the number of network layers can effectively balance the computational load and prediction accuracy, and maintain the real-time performance and stability of prediction.

[0079] Step 2.3: Implementation of the dynamic attention mechanism: In this step, a dynamic attention mechanism is designed to adaptively adjust the prediction weights of the three different objectives of flood control, power generation, and ecology. According to the weights of the flood control reward term, power generation reward term, and ecological reward term Update the target weights , and the update method is:

[0080] ;

[0081] ;

[0082] ;

[0083] After the above dynamic attention weights are adjusted in real time, they are used as the target weighting coefficients in the next step of reinforcement learning decision-making to ensure that the system can accurately adapt to the real-time needs of different scheduling objectives according to the real-time state;

[0084] The dynamic attention mechanism adopts a target weight vector with trainable parameters to calculate the attention weights of each target in real time using the Softmax function.

[0085] After the above dynamic attention weights are adjusted in real time, they are used as the target weighting coefficients in the next step of reinforcement learning decision-making to ensure that the system can accurately adapt to the real-time needs of different scheduling objectives according to the real-time state.

[0086] The dynamic attention mechanism adopts a target weight vector with trainable parameters to calculate the attention weights of each target in real time using the Softmax function.

[0087] Step 3: Implementation of the reinforcement learning multi-objective collaborative optimization module.

[0088] Step 3.1: Define the state space of reinforcement learning for multi-objective optimization. The state space includes meteorological data, the predicted water level sequence output by the adaptive network, the inflow sequence, and the outflow sequence. As Figure 3 shown, construct the state space of reinforcement learning. Based on the prediction results output by the dynamic Transformer network in Step 2, define the state vector of the reinforcement learning module as S t . Specifically, the state vector contains the predicted water level sequence for the next 3 time steps output by the Transformer network , the predicted inflow sequence , and the predicted outflow sequence C out,t = [C out,t+1 , C out,t+2, C out,t+3 and the latest real-time meteorological data sequence M t , where the meteorological data includes precipitation, wind speed, temperature and humidity. After splicing the above state data, the state space dimension is controlled between 100 and 500 dimensions, and after standardization processing, it is input into the reinforcement learning network. Inflow sequence, outflow sequence

[0089] Step 3.2: Define the reinforcement learning action space. Action vector a t contains three variables: flood discharge, power generation flow and ecological flow, and the specific settings are as follows:

[0090] Flood discharge regulation range: 50 to 5000 m³ / s, step size 50 m³ / s;

[0091] Power generation flow regulation range: 100 to 2000 m³ / s, step size 20 m³ / s;

[0092] Ecological flow regulation range: 10 to 500 m³ / s, step size 10 m³ / s;

[0093] The discrete combination of the action space is generated by grid search, and the state vector is mapped to the action space through a three-layer fully connected network for efficient action selection by the reinforcement learning model.

[0094] Step 3.3: Specific implementation of the reinforcement learning network. The reinforcement learning network uses a deep Q network, and the network structure is as follows:

[0095] Input layer: Input state vector S t ;

[0096] Hidden layer: Consists of 4 convolutional networks and 2 fully connected networks. The kernel sizes of the convolutional networks are [8×8, 4×4, 3×3, 3×3] in sequence, the convolutional strides are [4, 2, 1, 1] respectively, and the output feature dimensions are 128, 128, 64 and 64 respectively; then through two fully connected layers, the first layer outputs a 128-dimensional feature vector, using the ReLU activation function, and the second layer outputs the Q value of the action space;

[0097] Output layer: Q value of each action combination.

[0098] The optimization objective of the reinforcement learning network is to maximize the cumulative discounted reward R t , and the reward function is defined as the result of weighting three objectives (flood control, power generation, ecology), and the specific formula is:

[0099] ;

[0100] Among them, the reward weights α 防洪 , α 发电 , α生态 The calculation results of the dynamic attention mechanism from Step 2.3 are updated in real time. Specifically, the definitions of each target reward are as follows:

[0101] Flood control reward item R 防洪 Negative feedback is given according to the degree to which the predicted water level exceeds the flood control limit water level, and it is defined as: . Among them, H t is the predicted water level, and H 限 is the flood control limit water level.

[0102] Power generation reward item R 发电 It is calculated based on the economic benefits generated by the current power generation flow rate and is defined as: . Among them, k 发电 is the economic benefit coefficient of power generation per unit flow rate, and Q 发电,t is the power generation flow rate;

[0103] Ecological reward item R 生态 It is calculated based on the degree to which the downstream ecological flow rate deviates from the ideal ecological flow rate: . In the formula, Q 生态,t is the current ecological flow rate, and Q 生态,理想 is the ideal ecological flow rate.

[0104] During the training process, the parameters of the dynamic Transformer network are frozen, only the layer of output weights is activated, and it is jointly trained with the reinforcement learning network.

[0105] Through continuous iterative optimization, the reinforcement learning network finally converges to the optimal strategy for multi-objective collaborative optimization and outputs the optimal combination plan for each flow rate.

[0106] The network update learning rate of the reinforcement learning network is between 0.001 and 0.005; the dynamic range of the weight coefficients of each target in the reward function is 0.4 - 0.8 for flood control, 0.1 - 0.5 for power generation, and 0.1 - 0.3 for ecology, and the target weights are updated every 30 minutes.

[0107] Step 4: Implementation of the real-time reservoir operation decision-making module.

[0108] In this embodiment, the real-time reservoir operation decision-making module determines the opening degrees of each flood discharge gate of the reservoir and the output plan of the generator sets in real time through the optimal strategy results output by the reinforcement learning multi-objective collaborative optimization module, and forms specific execution instructions. Specifically, the flood discharge gate opening control instruction G t is obtained by real-time calculation through the gate flow rate - opening relationship formula from the flood discharge flow rate output by the reinforcement learning module:

[0109]

[0110] Among them, is an empirical relationship function between the flood discharge and the opening obtained from the actual measurement of the gate hydraulic characteristics.

[0111] The output adjustment command of the generator set is based on the generated flow rate output by reinforcement learning to determine the number of units started and stopped and the load distribution of each unit in real time:

[0112]

[0113] Among them, is the generated power of the i-th unit, ρ is the water density, g is the acceleration due to gravity, H is the water head, is the real-time water head, is the unit efficiency, is the number of units started.

[0114] On this basis, check whether the total discharge of "Q 泄洪,t +Q 发电,t " is not less than Q 生态,t . If it is satisfied, it is considered that the ecological water demand has been included and no additional scheduling is required; if it is insufficient, the difference Q 补生态,t =Q 生态,t- (Q 泄洪,t +Q 发电,t ) is used to supplement water through the ecological special sluice or low-load units to ensure that the downstream ecological flow reaches the standard, so as to realize the supplement of ecological flow.

[0115] Under extreme rainfall conditions (such as the rainfall intensity in the reservoir area exceeding 30 mm / h), the system automatically increases the flood control target weight α 防洪 to 0.85, and the reinforcement learning network outputs a high-intensity flood discharge strategy in real time to quickly respond to rainstorm events, so as to ensure that the water level in the reservoir area is always strictly controlled below the flood control limit water level and avoid the risk of flood disasters.

[0116] The response delay of the scheduling decision is controlled within 1 minute; under extreme rainfall conditions, the system automatically increases the flood control target weight to between 0.7 and 0.9, and adjusts the flood discharge at the fastest speed to ensure that the water level in the reservoir area is always controlled below the flood control limit water level.

[0117] Step 5: Implement the rolling optimization and feedback update mechanism.

[0118] Step 5.1: Real-time evaluation of prediction error and dynamic network update. The system automatically calculates the prediction error once every hour , and calculate the upper threshold and the lower threshold respectively according to step S2.2, and trigger the update process of the Transformer adaptive dynamic network structure and parameters according to the adjustment rule to reduce the subsequent prediction error. The network update is implemented by the gradient descent algorithm with a learning rate of 0.001, and it is immediately put back into real-time prediction work after the parameter adjustment is completed.

[0119] Step 5.2: Regular fine-tuning of the reinforcement learning strategy. At the end of each day, the system automatically counts the multi-objective optimization results of the reinforcement learning decisions, including the reduction of flood control risks, power generation benefits, and the guarantee of ecological flow, and conducts a comparative analysis with the objectives of the previous cycle. At the end of each week, the parameters of the reinforcement learning network are fine-tuned based on the summarized data. The experience replay strategy is adopted during the fine-tuning process, storing the state-action-reward data of the most recent week, and updating the parameters of the reinforcement learning network by randomly sampling in batches to ensure the continuous optimization and generalization performance of the network.

[0120] Step 5.3: Establishment and regular optimization of the long-term error database. The system establishes a historical database of long-term operation errors, and regularly conducts statistical analysis on the historical trends of prediction errors and decision-making errors every month. According to the historical error trends, the prediction error threshold (initially 0.05m, adjusted within the range of ±0.01m according to the trend) and the dynamic adjustment range of the reinforcement learning target weights α 防洪 , α 发电 , α 生态 are dynamically adjusted, and the specific adjustment range is controlled within ±10% to ensure the stability and adaptability during the long-term operation process.

[0121] Step 6: Deployment and implementation of the intelligent collaborative scheduling platform.

[0122] Such as Figure 4As shown in the figure, the implementation of the intelligent collaborative scheduling platform is achieved through an integrated intelligent management platform deployed on high-performance servers. This platform has functional modules such as real-time data monitoring, automatic execution of predictive decisions, real-time alarm, and historical data analysis. The platform adopts a distributed architecture and is deployed on Intel Xeon high-performance servers (CPU with 32 cores and 128GB of memory), with disaster recovery backup and remote Web access interfaces. The real-time data monitoring module receives and displays in real time the sensing data such as water level and meteorology; the adaptive network prediction module updates the prediction results every 10 minutes; the reinforcement learning decision module outputs optimized scheduling instructions in real time and automatically transmits them to on-site execution devices; the automatic scheduling execution module completes the real-time execution of the gate opening and power generation load. The response time of the system alarm mechanism is controlled within 30 seconds, which can automatically detect events such as water level exceeding the limit and extreme meteorological anomalies, trigger real-time alarms, and send them to management personnel via text messages and emails. The platform visualization interface includes the real-time reservoir status (such as the real-time water level curve and gate opening indication diagram), meteorological trend chart, multi-objective optimization decision-making chart, and supports the rapid query and analysis of historical data, providing comprehensive technical support for the comprehensive management of the reservoir in the form of charts and reports.

[0123] Step 7: Implement the long-term operation status monitoring and performance evaluation of the system.

[0124] To ensure the long-term stability and self-adaptability of the system of the present invention, a perfect long-term operation status monitoring and performance evaluation mechanism is established to regularly evaluate the comprehensive performance of the system. A comprehensive evaluation report on the system operation is generated quarterly, and the evaluation indicators include:

[0125] Flood control risk reduction rate:

[0126] ;

[0127] Power generation benefit improvement rate:

[0128] ;

[0129] Ecological flow guarantee rate:

[0130] ;

[0131] The evaluation error of each indicator is controlled within ±5%. According to the quarterly evaluation results, if there is a significant downward trend (the decline exceeds 10%) in any indicator, the update and structural adjustment process of the adaptive dynamic network structure and reinforcement learning network parameters will be automatically triggered to restore the system performance. In addition, a system operation log and abnormal event recording mechanism is established to long-term store and manage the changes in scheduling strategies, prediction anomalies, parameter adjustment records, etc., forming a complete operation database for annual technical upgrade and system maintenance decisions.

[0132] Through real-time data fusion preprocessing technology, adaptive dynamic network structure design, and multi-objective reinforcement learning decision-making mechanism, the present invention comprehensively improves the collaborative management efficiency of multiple objectives such as reservoir flood control safety, power generation benefits, and ecological protection. Especially under extreme weather conditions, the present invention can quickly respond and dynamically adjust the scheduling objective weights, effectively reducing the decision-making delay and risk in the traditional reservoir scheduling method. Through the rolling update and feedback optimization mechanism of the dynamic network structure and reinforcement learning decision-making, the present invention can maintain high adaptability and stability for a long time, meeting the strict requirements of real-time and accuracy in actual reservoir management. This method can be widely applied to the intelligent and refined management of large and medium-sized reservoirs, significantly improving the reservoir scheduling decision-making level and the safety operation guarantee ability.

[0133] The present invention provides a reservoir scheduling method based on deep learning adaptive dynamic network and reinforcement learning. Through real-time water level sensors, meteorological data acquisition systems, and high-resolution satellite remote sensing equipment, multi-source data is fused and accurately preprocessed in real time, effectively improving the accuracy and integrity of reservoir real-time state data. By designing an adaptive dynamic network and a dynamic attention mechanism based on the Transformer structure, the present invention realizes short-term accurate prediction of reservoir water level, inflow, and outflow, and can automatically and dynamically adjust the network structure according to the prediction error to ensure a good balance between prediction accuracy and calculation efficiency. At the same time, the present invention combines the deep reinforcement learning method to construct a multi-objective collaborative optimization decision-making module, and based on the dynamic prediction results, it outputs the optimal combination strategy of flood discharge, power generation, and ecological flow in real time, realizing the real-time dynamic intelligent collaborative regulation of multiple objectives of the reservoir. In the real-time scheduling execution link, the present invention can quickly convert the optimization strategy into accurate execution instructions for gate opening and generator set output, and quickly respond under extreme weather conditions to ensure flood control safety.

[0134] The present invention proposes a complete and innovative rolling optimization feedback update mechanism. By real-time evaluating the prediction error and the effect of reinforcement learning decision-making, it automatically triggers the fine-tuning of network parameters and policy updates, effectively enhancing the stability and adaptability during long-term operation. The deployment of the intelligent collaborative scheduling platform realizes real-time data monitoring, automatic execution of prediction and decision-making, and rapid alarm for abnormal events, effectively improving the automation, refinement, and intelligence levels of reservoir management. In addition, the present invention further establishes a monitoring and evaluation mechanism for the long-term operation state and performance of the system, which can regularly evaluate the achievement of flood control safety, power generation benefits, and ecological protection goals, and adaptively optimize the network structure and parameters according to the evaluation results, enabling the long-term operation performance of the system to remain in an optimal state. Experimental verification in an actual reservoir environment shows that the implementation of the present invention can significantly reduce the flood control risk of the reservoir, improve the economic benefits of power generation, and effectively guarantee the ecological flow, with the outstanding advantages of rapid real-time response, strong multi-objective collaborative optimization ability, and high long-term operation stability, and can effectively meet the management requirements of modern reservoirs in complex meteorological and hydrological environments.

[0135] The reservoir scheduling method based on the deep learning adaptive dynamic network proposed by the present invention solves.

[0136] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A reservoir operation method based on a deep learning adaptive dynamic network and reinforcement learning, characterized in that, Including: Obtain the real-time water level data of the reservoir, meteorological data, and radar image data of the regional cloud layer and surface characteristics, and perform data fusion on the three; Input the fused data into a preset adaptive dynamic Transformer network model to obtain a prediction output, where the prediction output includes a reservoir water level sequence, an incoming flow sequence, and an outflow sequence; Construct a reinforcement learning state space, which includes the predicted reservoir water level sequence, incoming flow sequence, outflow sequence, and the latest real-time meteorological data sequence at a preset time step in the future of the Transformer network model; Input the reinforcement learning state space into a preset reinforcement learning network. The action space of the reinforcement learning network includes the flood discharge volume, power generation flow rate, and ecological flow rate. The optimization objective of the reinforcement learning network is to maximize the cumulative discounted reward, and the reward function is a weighted function of the flood control reward term, power generation reward term, and ecological reward term. Through continuous iterative optimization, the reinforcement learning network finally converges to output the optimal combination plan of the flood discharge volume, power generation flow rate, and ecological flow rate.

2. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1, characterized in that Before data fusion, preprocess the real-time water level data and meteorological data. The preprocessing steps include: Use the Kriging interpolation algorithm based on spatial distance weights to perform spatial interpolation on the real-time water level data, meteorological data, and radar image data of the regional cloud layer and surface characteristics to obtain continuous data covering the entire reservoir area; Detect outliers in the continuous data and remove them in real time; Perform smoothing and denoising processing on the data after removing outliers.

3. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1, characterized in that, When the error between the prediction output of the adaptive dynamic Transformer network model and the actual observation result exceeds the preset upper threshold each time, automatically add a preset number of encoding layers to the adaptive dynamic Transformer network model. When the result error is lower than the preset lower threshold, automatically reduce a preset number of encoding layers from the adaptive dynamic Transformer network model. When the error between the prediction output of the adaptive dynamic Transformer network model and the actual observation result is between the lower threshold and the upper threshold, the adaptive dynamic Transformer network model remains unchanged.

4. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 3, characterized in that The result error is the weighted sum of the mean square error of the predicted and actual reservoir water levels of the reservoir water level, the mean square error of the predicted and actual incoming flow rates of the incoming flow, and the mean square error of the predicted and actual outflow rates of the outflow.

5. The reservoir operation method based on the deep learning adaptive dynamic network and reinforcement learning according to claim 1, characterized in that The adaptive dynamic Transformer network model outputs the weights of the flood control reward term, power generation reward term, and ecological reward term in the reward function based on the dynamic attention mechanism each time.

6. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 5, characterized in that, According to the weights of flood control reward items, power generation reward items and ecological reward items Update the target weights , and the update method is as follows: ; ; ; The dynamic attention mechanism uses a target weight vector with trainable parameters to calculate the attention weight of each target in real time using the Softmax function.

7. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 6, wherein The formula of the reward function is as follows: ; ; ; ; R t For the cumulative discount reward, α 防洪 is the weight coefficient of the flood control reward item, α 发电 is the weight coefficient of the power generation reward item, α 生态 is the weight coefficient of the ecological reward item, H t is the predicted water level, H 限 is the flood control limit water level, k 发电 is the economic benefit coefficient of power generation per unit flow, Q 发电,t is the power generation flow, Q 生态,t is the current ecological flow, Q 生态,理想 is the ideal ecological flow, R 防洪 is the flood control reward item, R 发电 is the power generation reward item, R 生态 is the ecological reward item, and t is the time step.

8. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 1, characterized in that Determine the opening of each flood discharge gate of the reservoir and the output plan of the generator set according to the optimal combination plan of the flood discharge volume, power generation flow rate, and ecological flow rate output, and form specific execution instructions, which specifically include: Calculate the flood discharge gate opening in real time according to the output flood discharge volume and the gate flow-opening relationship formula; Determine the number of generating units to start and stop and the load distribution of each generating unit in real time according to the generated power flow output; Construct an execution instruction based on the opening of flood discharge gates, the number of generating units to start and stop, and the load distribution of each generating unit; Judge whether the total discharge of "flood discharge + power generation" is not lower than the ecological flow. If so, it is considered that the ecological water demand has been included and no additional scheduling is required. If not, make up the water through the ecological special sluice opening or low-load generating units according to the difference, where the difference is the difference between the ecological flow and the total discharge of "flood discharge + power generation".

9. The reservoir operation method based on deep learning adaptive dynamic network and reinforcement learning according to claim 8, characterized in that Construct a historical database of long-term running root mean square error, regularly conduct statistical analysis on the historical trends of prediction errors and decision-making errors, and dynamically adjust the preset error threshold based on the historical error trends; 10. The reservoir operation method based on the deep learning adaptive dynamic network and reinforcement learning according to claim 1, wherein, Obtain the real-time water level data of the reservoir, meteorological data, and radar image data of regional cloud layers and surface characteristics, including: Obtain the real-time water level data monitored by a wave-type water level sensor, and the wave-type water level sensor is deployed at key cross-sections of the reservoir area and the main incoming river channels; Obtain the meteorological data of the meteorological data acquisition system, and the meteorological data acquisition system is deployed within the reservoir area. The meteorological data acquisition system includes a precipitation sensor, a wind speed sensor, and a temperature and humidity sensor; Obtain the radar image data of regional cloud layers and surface characteristics. Use a high-resolution radar remote sensing satellite with a spatial resolution of 30m to conduct satellite remote sensing observations on the reservoir and the upstream area of the basin, and obtain the radar image data of regional cloud layers and surface characteristics.

Citation Information

Patent Citations

  • Train autonomous scheduling deep reinforcement learning method and module

    CN111369181A

  • Reservoir group joint optimization scheduling method based on MADDPG reinforcement learning

    CN115952958A

  • Flood warning method and system integrating meteorological and hydrological sensitivity

    CN117575873A

  • Method and system for predicting power generation capacity of hydropower station

    CN119721368A

  • Predicting gas lift equipment failure with deep learning techniques

    US20250075602A1

Cited By

  • Reservoir flood season water level dynamic multi-objective optimization control system and method

    CN120745946A

  • A reservoir flood season water level dynamic multi-objective optimization control system and method

    CN120745946B

  • Reservoir scheduling optimization method and system based on aerospace big data

    CN121766733A

  • Intelligent water conservancy centralized control center system based on intelligent on-duty body cluster and scheduling method

    CN121996392A