Reservoir scheduling method and device based on multi-objective reinforcement learning
Through the reservoir scheduling method of multi-objective reinforcement learning, combined with regression loss and state loss, the flexibility and adaptability problems of traditional reservoir scheduling methods are solved, accurate and intelligent decision-making of reservoir scheduling is achieved, and multi-objective optimization is adapted to complex environments.
Patent Information
- Application Number
- CN202510524209.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional reservoir scheduling methods are difficult to make real-time and accurate decisions when facing complex and changing climatic conditions and sudden natural factors, and lack flexibility and adaptability. In the reservoir scheduling, deep learning models have problems such as data quality affecting prediction capabilities, multi-objective optimization problems, and lack of interpretability in decision-making processes.
The reservoir scheduling method based on multi-objective reinforcement learning is adopted. By obtaining real data, a multi-objective deep reinforcement learning network is built, and parameter updates and iterative training are performed to output real-time optimized reservoir scheduling strategies.
It realizes accurate and intelligent decision-making support for reservoir scheduling, has the ability to quickly iterate and optimize, can adapt to different types of reservoir and resource scheduling problems, has efficient dynamic reconstruction capabilities and strong scalability.
Smart Images

Figure CN120355170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reservoir operation, and particularly to a reservoir operation method and device based on multi-objective reinforcement learning. Background Art
[0002] Reservoir operation is a core issue in water resources management, aiming to rationally allocate the water resources in the reservoir to meet various demands, including water supply, power generation, irrigation, ecological protection, and flood control. Traditional reservoir operation methods mainly rely on expert experience and rule-based operation strategies, often using rule algorithms and numerical models to perform operations such as water level control and flow regulation. Although these methods can provide certain operation benefits in some cases, they are often difficult to make real-time and accurate decisions in the face of complex and changing climatic conditions, sudden floods, extreme droughts and other natural factors. In addition, the traditional methods respond relatively slowly to changes in the internal and external environment of the reservoir, lacking flexibility and self-adaptability. Therefore, with the diversification of water resources management requirements, the limitations of traditional operation methods are becoming increasingly prominent.
[0003] To solve the problems existing in traditional operation methods, the prior art adopts a reservoir operation method based on deep learning to achieve adaptive operation in complex and dynamic environments. However, although the reservoir operation method based on deep learning has shown great potential, it still faces many challenges and dilemmas in practical applications. First, the training of deep learning models requires a large amount of high-quality data, and the historical data involved in reservoir operation often has missing values, noise and inconsistencies, and the data quality directly affects the prediction ability of the model. Second, the training process of deep learning models is complex and time-consuming. Especially when facing multi-objective optimization problems, how to design a suitable loss function to balance the weights between different objectives is still a difficult problem. The objectives in reservoir operation (such as power generation, water supply, flood control) often conflict with each other. How to achieve a reasonable balance of objectives while ensuring the efficiency of the model is a major difficulty in applying deep learning to reservoir operation. In addition, the "black box" characteristic of deep learning models makes their decision-making process lack interpretability, and it is difficult to explain the decision-making basis of the model in practical applications, resulting in doubts about the credibility of the model when decision-makers use it. Finally, the deployment and real-time response ability of deep learning in reservoir operation also face challenges. Especially how to ensure the generalization ability and real-time performance of the model in complex environments still requires further research and optimization. Summary of the Invention
[0004] The present invention proposes a reservoir operation method and device based on multi-objective reinforcement learning to solve the problems of inaccurate prediction and unbalanced operation of existing reservoir operation methods.
[0005] The present invention realizes the above object through the following technical solutions:
[0006] A reservoir operation method based on multi-objective reinforcement learning of the present invention includes:
[0007] Step 1: Obtain information and real data of reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir operation strategies, and external environment data. The reservoir operation strategies include flood discharge and water storage. Among them, the reservoir characteristic parameters include reservoir capacity, reservoir area, normal storage level, dead storage level, characteristic water levels, and flood discharge capacity; historical operation information of various types includes historical water levels, historical water storage, historical inflow, and historical outflow; external environment data includes precipitation, temperature, incoming floods, flood control, power generation, water supply and irrigation demands, and runoff data.
[0008] Step 2: Based on a reservoir operation simulation software, build a reservoir operation model, input the information into the reservoir operation model, and output simulated data of reservoir operation. Both the real data and the simulated data of reservoir operation include changes in reservoir water level, changes in outflow, changes in reservoir water storage, power generation, water storage, utilization rate of flood resources, and water supply guarantee rate.
[0009] Step 3: Preprocess the real data and the simulated data to obtain preprocessed data, and construct a data set according to the preprocessed data. The data set includes several groups of training space state data and corresponding reservoir operation strategies. The training space state data includes current water level, precipitation, inflow, outflow, temperature, historical water level changes, historical precipitation changes, power generation demand, and water use demand.
[0010] Step 4: Input a group of the training space state data into a preset multi-objective deep reinforcement learning network and output a predicted reservoir operation strategy.
[0011] Step 5: Input the reservoir structure characteristic parameters, reservoir historical operation data, and the predicted reservoir operation strategy into the reservoir operation model and output predicted space state data.
[0012] Step 6: Solve the regression distance between the predicted flood discharge and water storage and the real flood discharge and water storage to obtain a regression loss. Calculate the state loss through a reward function based on the power generation, water supply, and current water level in the predicted space state data and the preset water supply demand, power generation demand, normal storage level, and safety water level. Construct a loss function according to the regression loss and the state loss.
[0013] Step 7: Update the parameters of the multi-objective deep reinforcement learning network based on the loss function in the backpropagation manner. Use the predicted spatial state data as the new training spatial state data, input it into the multi-objective deep reinforcement learning network with updated parameters, and iterate and update according to Steps 5 to 7 until the regression loss is zero, obtaining the updated target deep reinforcement learning network;
[0014] Step 8: Input the training spatial state data of other groups in the several groups of training spatial state data into the target deep reinforcement learning network iteratively trained and updated according to Steps 4 to 7 to obtain a trained multi-objective deep reinforcement learning network;
[0015] Step 9: Input the real-time collected spatial state data into the trained multi-objective deep reinforcement learning network and output the real-time optimized reservoir operation strategy.
[0016] Furthermore, preprocess the real data and the simulated data, including:
[0017] Remove outliers, fill in missing values, and perform standardization processing on the real data and the simulated data, and classify the data into three levels: general, important, and very important;
[0018] The method for filling in missing values is the linear interpolation method, and the formula is:
[0019] y is the interpolation result at the missing value position, is the abscissa of the known data point, is the ordinate of the known data point, and x is the abscissa of the missing value position;
[0020] In the standardization process, the maximum-minimum normalization method is used for water level, inflow, outflow, temperature, flood discharge, and water storage, and the Z-score normalization method is used for precipitation, historical water level change, and historical precipitation change.
[0021] Further, the multi-objective deep reinforcement learning network has a three-input and two-output network structure. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with dimensions of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The multi-objective deep reinforcement learning network includes a feature compression module and a feature extraction module. The feature compression module is used to extract features from the input feature matrix using a convolutional kernel, then superimpose the extraction result and the input data and perform normalization processing, and finally perform pooling processing. The feature extraction module is used to first extract features from the input feature matrix using a convolutional kernel through the feature extraction module, and finally superimpose the extraction result and the input data and perform normalization processing. The multi-objective deep reinforcement learning network is used to process the historical water level change vector and the historical precipitation change vector respectively using a feature compression module with 512, 256, 128, and 64 convolutional kernels to obtain two compressed feature vectors of 1*8*64, and process the key parameter vector using a feature extraction module with 512, 256, 128, and 64 convolutional kernels to obtain a feature vector of 1*8*64. Concatenate this feature vector with the compressed feature vectors respectively to obtain two feature vectors of 1*8*128, then process the two feature vectors respectively using a feature extraction module with 64 and 32 convolutional kernels to obtain two feature vectors of 1*8*32, concatenate them to obtain a feature vector of 1*8*64, perform high-order feature semantic extraction using a feature extraction module with 64 convolutional kernels respectively, after flattening, use a fully connected layer of 512 for feature compression, and finally use a fully connected layer with a dimension of 1 to complete the final output.
[0022] Further, the loss function is:
[0023] Where, and are the weight values of the flood discharge volume and the water storage volume in the total loss function respectively, with a range of 0-1; is the normalized value of the predicted flood discharge volume and the normalized value of the actual flood discharge volume; is the normalized value of the predicted water storage volume and the normalized value of the actual water storage volume; N is the total number of training space states in the training data; w1 is the weight parameter of the power generation reward, w2 is the weight parameter of the water supply reward, and w3 is the weight parameter of the flood control reward; is the normalized value of the actual power generation volume and the normalized value of the power generation volume demanded by the power grid; is the electricity price coefficient, with a range of 0-1; are the actual water supply volume and the minimum / maximum water supply demand respectively; is the current water level and the safety water level; is the flood control penalty coefficient, which is used to control the intensity of flood control rewards and ranges from 0 to 1.
[0024] The present invention also provides a reservoir scheduling device based on multi-objective reinforcement learning, comprising:
[0025] an acquisition module, which is used to acquire information and real data of reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir scheduling strategies, and external environment data. The reservoir scheduling strategies include flood discharge volume and water storage volume;
[0026] a first output module, which is used to model a reservoir scheduling model based on reservoir scheduling simulation software, input the information into the reservoir scheduling model, and output simulated data of reservoir operation. Both the real data and the simulated data of reservoir operation include reservoir water level changes, outflow discharge changes, reservoir water storage volume changes, power generation, water storage volume, flood water resource utilization rate, and water supply guarantee rate;
[0027] a preprocessing module, which is used to preprocess the real data and the simulated data to obtain preprocessed data, and construct a data set according to the preprocessed data. The data set includes several groups of training space state data and corresponding reservoir scheduling strategies. The training space state data includes the current water level, precipitation, inflow discharge, outflow discharge, temperature, historical water level changes, historical precipitation changes, power generation demand, and water use demand;
[0028] a second output module, which is used to input a group of the training space state data into a preset multi-objective deep reinforcement learning network and output a predicted reservoir scheduling strategy;
[0029] a third output module, which is used to input the reservoir structure characteristic parameters, reservoir historical operation data, and the predicted reservoir scheduling strategy into the reservoir scheduling model and output predicted space state data;
[0030] a construction module, which is used to solve the regression distance according to the predicted flood discharge volume and water storage volume and the real flood discharge volume and water storage volume to obtain a regression loss, calculate a state loss according to the power generation, water supply, current water level in the predicted space state data and the preset water supply demand, power generation demand, normal storage water level, and safety water level through a reward function, and construct a loss function according to the regression loss and the state loss;
[0031] An update module, which is used to update the parameters of the multi-objective deep reinforcement learning network based on the loss function in a backpropagation manner, take the predicted spatial state data as new training spatial state data, input it into the multi-objective deep reinforcement learning network with updated parameters, and iterate and update according to steps five to seven until the regression loss is zero, to obtain an updated target deep reinforcement learning network;
[0032] A training module, which is used to iteratively train the updated target deep reinforcement learning network according to the input of the training spatial state data of other groups in the several groups of training spatial state data, to obtain a trained multi-objective deep reinforcement learning network;
[0033] An optimization strategy module, which is used to input the real-time collected spatial state data into the trained multi-objective deep reinforcement learning network and output a real-time optimized reservoir operation strategy.
[0034] Furthermore, preprocess the real data and the simulation data, including:
[0035] Remove outliers, fill in missing values, and standardize the real data and the simulation data, and classify the data into three levels: general, important, and very important;
[0036] The method for filling in missing values is the linear interpolation method, and the formula is:
[0037] ;
[0038] y is the interpolation result at the missing value position, x1 and x2 are the abscissas of known data points, y1 and y2 are the ordinates of known data points, and x is the abscissa of the missing value position;
[0039] In the standardization process, the maximum-minimum normalization method is used for water level, inflow, outflow, temperature, flood discharge, and water storage, and the Z-score normalization method is used for precipitation, historical water level change, and historical precipitation change.
[0040] Further, the multi-objective deep reinforcement learning network has a three-input and two-output network structure. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with sizes of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The multi-objective deep reinforcement learning network includes a feature compression module and a feature extraction module. The feature compression module is used to extract features from the input feature matrix using a convolution kernel, then superimpose the extraction result and the input data in a matrix and perform normalization processing, and finally perform pooling processing. The feature extraction module is used to first extract features from the input feature matrix using a convolution kernel through the feature extraction module, and finally superimpose the extraction result and the input data in a matrix and perform normalization processing. The multi-objective deep reinforcement learning network is used to process the historical water level change vector and the historical precipitation change vector respectively using a feature compression module with 512, 256, 128, and 64 convolution kernels to obtain two compressed feature vectors of 1*8*64, and process the key parameter vector using a feature extraction module with 512, 256, 128, and 64 convolution kernels to obtain a feature vector of 1*8*64. Concatenate this feature vector with the compressed feature vectors respectively to obtain two feature vectors of 1*8*128, then process the two feature vectors respectively using a feature extraction module with 64 and 32 convolution kernels to obtain two feature vectors of 1*8*32, concatenate them to obtain a feature vector of 1*8*64, perform high-order feature semantic extraction using a feature extraction module with 64 convolution kernels respectively, flatten them, perform feature compression using a fully connected layer with 512, and finally complete the final output using a fully connected layer with a dimension of 1.
[0041] Further, the loss function is:
[0042] ;
[0043] Among them, and are the weight values of the flood discharge volume and the water storage volume in the total loss function respectively, with a range of 0-1; is the normalized value of the predicted flood discharge volume and the normalized value of the actual flood discharge volume; is the normalized value of the predicted water storage volume and the normalized value of the actual water storage volume; N is the total number of training space states in the training data; is the weight parameter of the power generation reward, is the weight parameter of the water supply reward, is the weight parameter of the flood control reward; is the normalized value of the actual power generation volume and the normalized value of the power generation volume demanded by the power grid; is the electricity price coefficient, with a range of 0 - 1; are the actual water supply volume and the minimum / maximum water supply demand respectively; are the current water level and the safety water level; is the flood control penalty coefficient, used to control the intensity of flood control rewards, with a range of 0 - 1.
[0044] The beneficial effects of the present invention are as follows:
[0045] A reservoir scheduling method based on multi - objective reinforcement learning proposed by the present invention combines regression loss and state loss, promotes the rapid iteration and optimization of the reinforcement learning model in reservoir scheduling, makes full use of data - driven modeling and decision - making methods, realizes precise and intelligent decision - making support in reservoir scheduling, has strong scalability and generality, can be flexibly applied in various reservoir scheduling scenarios, and can adapt to different types of reservoirs and resource scheduling problems. Description of the Drawings
[0046] Figure 1 is the model training process of a reservoir scheduling method based on multi - objective reinforcement learning of the present invention.
[0047] Figure 2 is the model inference process of a reservoir scheduling method based on multi - objective reinforcement learning of the present invention.
[0048] Figure 3 is a construction method of multi - objective reinforcement learning of the present invention.
[0049] Figure 4 is a schematic diagram of a multi - objective deep reinforcement learning network based on fast optimization of the present invention. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0051] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0052] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0053] In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0054] A reservoir operation scheduling method based on multi-objective reinforcement learning of the present invention includes:
[0055] Step 1: Obtain information and real data of reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir operation scheduling strategies, and external environment data. The reservoir operation scheduling strategies include flood discharge volume and water storage volume.
[0056] Step 2: Based on reservoir operation simulation software, build a reservoir operation scheduling model. Input the information into the reservoir operation scheduling model and output simulated data of reservoir operation. Both the real data and the simulated data of reservoir operation include reservoir water level changes, outflow discharge changes, reservoir water storage volume changes, power generation, water storage volume, flood water resource utilization rate, and water supply guarantee rate.
[0057] Step 3: Preprocess the real data and the simulated data to obtain preprocessed data, and construct a data set according to the preprocessed data. The data set includes several groups of training space state data and corresponding reservoir operation scheduling strategies. The training space state data includes current water level, precipitation, inflow discharge, outflow discharge, temperature, historical water level changes, historical precipitation changes, power generation demand, and water use demand.
[0058] Step 4: Input a group of the training space state data into a preset multi-objective deep reinforcement learning network and output a predicted reservoir operation scheduling strategy.
[0059] Step 5: Input the reservoir structure characteristic parameters, reservoir historical operation data, and the predicted reservoir operation scheduling strategy into the reservoir operation scheduling model and output predicted space state data.
[0060] Step 6: Solve the regression distance according to the predicted flood discharge volume and water storage volume and the real flood discharge volume and water storage volume to obtain a regression loss. Calculate a state loss according to the power generation, water supply, and current water level in the predicted space state data and preset water supply demand and power generation demand through a reward function. Construct a loss function according to the regression loss and the state loss.
[0061] Step 7: Update the parameters of the multi-objective deep reinforcement learning network based on the loss function in the reverse propagation manner. Use the predicted spatial state data as the new training spatial state data, input it into the multi-objective deep reinforcement learning network with updated parameters, and iterate and update according to Steps 5 to 7 until the regression loss is zero, obtaining the updated target deep reinforcement learning network;
[0062] Step 8: Input the training spatial state data of other groups in the several groups of training spatial state data into the target deep reinforcement learning network iteratively trained and updated according to Steps 4 to 7, obtaining a trained multi-objective deep reinforcement learning network;
[0063] Step 9: Input the real-time collected spatial state data into the trained multi-objective deep reinforcement learning network, and output the real-time optimized reservoir operation strategy.
[0064] In some embodiments, preprocess the real data and the simulated data, including:
[0065] Remove outliers, fill in missing values, and perform standardization processing on the real data and the simulated data, and classify the data into three levels: general, important, and very important;
[0066] The method for filling in missing values is the linear interpolation method, and the formula is:
[0067]
[0068] y is the interpolation result at the missing value position, is the abscissa of the known data point, is the ordinate of the known data point, and x is the abscissa of the missing value position;
[0069] In the standardization processing, the maximum-minimum normalization method is used for water level, inflow, outflow, temperature, flood discharge, and water storage, and the Z-score normalization method is used for precipitation, historical water level change, and historical precipitation change.
[0070] In some embodiments, the multi-objective deep reinforcement learning network has a three-input and two-output network structure. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with dimensions of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The multi-objective deep reinforcement learning network includes a feature compression module and a feature extraction module. The feature compression module is used to extract features from the input feature matrix using a convolutional kernel, then stack the extraction result and the input data for matrix addition and perform normalization processing, and finally perform pooling processing. The feature extraction module is used to first extract features from the input feature matrix using a convolutional kernel through the feature extraction module, and finally stack the extraction result and the input data for matrix addition and perform normalization processing. The multi-objective deep reinforcement learning network is used to process the historical water level change vector and the historical precipitation change vector using the feature compression module with 512, 256, 128, and 64 convolutional kernels respectively to obtain two compressed feature vectors of 1*8*64, process the key parameter vector using the feature extraction module with 512, 256, 128, and 64 convolutional kernels to obtain a feature vector of 1*8*64, splice the feature vector with the compressed feature vectors respectively to obtain two feature vectors of 1*8*128, then process the two feature vectors using the feature extraction module with 64 and 32 convolutional kernels respectively to obtain two feature vectors of 1*8*32, splice them to obtain a feature vector of 1*8*64, perform high-order feature semantic extraction using the feature extraction module with 64 convolutional kernels, flatten them, perform feature compression using a fully connected layer with 512, and finally complete the final output using a fully connected layer with a dimension of 1.
[0071] In some embodiments, the loss function is:
[0072]
[0073] Wherein, and are the weight values of the flood discharge volume and the water storage volume in the total loss function respectively, with a range of 0-1; is the normalized value of the predicted flood discharge volume and the normalized value of the actual flood discharge volume; is the normalized value of the predicted water storage volume and the normalized value of the actual water storage volume; N is the total number of training space states in the training data; w1 is the weight parameter of the power generation reward, w2 is the weight parameter of the water supply reward, and w3 is the weight parameter of the flood control reward; is the normalized value of the actual power generation volume and the normalized value of the power generation volume required by the power grid; is the electricity price coefficient, with a range of 0-1; They are the actual water supply volume, the minimum and maximum water demand respectively; They are the current water level and the safety water level; It is the flood control penalty coefficient, which is used to control the intensity of flood control rewards, and its range is 0 - 1. It represents the water supply guarantee rate. It represents the utilization rate of flood resources.
[0074] The present invention also provides a reservoir scheduling device based on multi - objective reinforcement learning, including:
[0075] An acquisition module, which is used to acquire information and real - time data of reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir scheduling strategies, and external environment data. The reservoir scheduling strategies include flood discharge volume and water storage volume.
[0076] A first output module, which is used to model a reservoir scheduling model based on reservoir scheduling simulation software, input the information into the reservoir scheduling model, and output simulated data of reservoir operation. Both the real - time data and the simulated data of reservoir operation include reservoir water level changes, outlet flow changes, reservoir water storage changes, power generation, water storage volume, utilization rate of flood resources, and water supply guarantee rate.
[0077] A pre - processing module, which is used to pre - process the real - time data and the simulated data to obtain pre - processed data, and construct a data set according to the pre - processed data. The data set includes several groups of training space state data and corresponding reservoir scheduling strategies. The training space state data includes the current water level, precipitation, inflow, outlet flow, temperature, historical water level changes, historical precipitation changes, power generation demand, and water use demand.
[0078] A second output module, which is used to input a group of the training space state data into a preset multi - objective deep reinforcement learning network and output a predicted reservoir scheduling strategy.
[0079] A third output module, which is used to input the reservoir structure characteristic parameters, reservoir historical operation data, and the predicted reservoir scheduling strategy into the reservoir scheduling model and output predicted space state data.
[0080] A construction module, which is used to solve the regression distance between the predicted flood discharge volume and water storage volume and the real flood discharge volume and water storage volume to obtain a regression loss, calculate a state loss according to the power generation, water supply volume, and current water level in the predicted space state data and the preset water supply demand and power generation demand through a reward function, and construct a loss function according to the regression loss and the state loss.
[0081] An update module, which is configured to update the parameters of the multi-objective deep reinforcement learning network based on the loss function in a backpropagation manner, use the predicted spatial state data as new training spatial state data, input it into the multi-objective deep reinforcement learning network with updated parameters, and iterate and update according to steps five to seven until the regression loss is zero, so as to obtain an updated target deep reinforcement learning network;
[0082] A training module, which is configured to iteratively train the updated target deep reinforcement learning network according to the input of the training spatial state data of other groups in the several groups of training spatial state data, so as to obtain a trained multi-objective deep reinforcement learning network;
[0083] An optimization strategy module, which is configured to input the real-time collected spatial state data into the trained multi-objective deep reinforcement learning network and output a real-time optimized reservoir operation strategy.
[0084] The model training process of a reservoir operation method based on multi-objective reinforcement learning according to the present invention is as Figure 1As shown in the figure, first, by obtaining various historical operation information of the reservoir, such as state information like water level, water storage volume, reservoir capacity, discharge capacity, and incoming water volume, and combining with external environmental data, such as the demands of flood control, power generation, water supply, and irrigation, the operation situation of the reservoir is comprehensively understood. Then, a reservoir operation simulation software is used for modeling, and reservoir characteristic parameters, historical runoff data, operation rules or strategies, etc. are input to simulate the operation state of the real reservoir, and simulation data such as changes in reservoir water level, outflow discharge, and reservoir water storage volume are obtained. At the same time, the operation effect evaluation is carried out, including power generation, water storage volume, utilization rate of flood resources, and water supply guarantee rate, etc. Then, the real and simulated data are cleaned and preprocessed, outliers are removed, missing values are filled, and standardization processing is carried out. In addition, according to the different importance of the data, it is divided into three levels: general, important, and very important, so as to give different weights during the training process. The processed data is organized into a data set. The input part includes the current water level, precipitation, incoming flow, outflow discharge, temperature, historical water level changes, historical precipitation changes, power generation demand, water use demand, etc., and the output part is the flood discharge volume and water storage volume. The data set contains multiple training space state data and corresponding operation methods. Then, a training space state data is input into a multi-objective reinforcement learning network based on fast optimization. The network outputs the execution actions of the flood discharge volume and water storage volume for the next step according to the current state. Information such as the reservoir characteristic parameters, historical runoff data, and the flood discharge volume and water storage volume output by the network are then input into the reservoir operation simulation software, and the software returns data such as the current water level, precipitation, incoming flow, outflow discharge, power generation, and water supply. The regression distance between the predicted flood discharge volume and water storage volume and the real values is solved to obtain the regression loss. At the same time, the losses between power generation, water supply, the current water level and water supply demand, and power generation demand are calculated according to the reward function. Finally, the sum of the regression loss and the state loss is obtained, and the parameters of the reinforcement learning model are updated through backpropagation. Next, the model training is continuously iterated until the regression loss function of the model approaches zero. When all the training data are completed, the training is stopped and the reinforcement learning model is saved.
[0085] Among them, the model inference process of the reservoir operation method based on multi-objective reinforcement learning of the present invention is as Figure 2As shown in the figure, the real reservoir data is preprocessed and organized into a dataset format, and the training space state data is input into a multi-objective reinforcement learning network based on fast optimization. The network outputs the execution actions of the next flood discharge volume and water storage volume according to the current state. Information such as the characteristic parameters of the reservoir, historical runoff data, and the flood discharge volume and water storage volume output by the network is then input into the reservoir operation simulation software, and the software returns data such as the current water level, precipitation, inflow, outflow, power generation, and water supply. The regression distance between the predicted flood discharge volume and water storage volume and the real values is solved to obtain the regression loss. At the same time, the losses between the power generation, water supply, current water level and water supply demand, and power generation demand are calculated according to the reward function. Finally, the sum of the regression loss and the state loss is obtained, and the parameters of the reinforcement learning model are updated through backpropagation. Then, the model training is continuously iterated by constructing the latest space state until the regression loss function of the model approaches zero. The flood discharge volume and water storage volume of each state in the process are recorded to obtain an optimal reservoir operation strategy method.
[0086] The construction method of the multi-objective reinforcement learning based on the present invention is as Figure 3 shown. The reinforcement learning space state of the reservoir operation includes the current water level, precipitation, inflow and outflow. At the same time, temperature, humidity, historical water level changes, historical precipitation changes, power generation demand, water use demand, and space state parameters are used as the inputs of the deep reinforcement learning model. The model outputs the flood discharge volume and water storage volume as the action strategy of the reinforcement learning. The reward evaluation of the power generation benefit, water supply guarantee, flood control safety, and water storage capacity is carried out through the influence brought by the action strategy, so as to update the parameters of the deep reinforcement learning and update the state space parameter information at the same time.
[0087] The schematic diagram of the multi-objective deep reinforcement learning network based on fast optimization designed by the present invention is as Figure 4As shown, its network structure is a three-input and two-output network structure. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with dimensions of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The network uses convolution, normalization, and pooling as the feature compression module (Block1). The compression module first uses a convolution kernel to extract features from the input feature matrix, and the extraction result is matrix-added to the input data and then normalized. Finally, pooling is performed. Convolution and normalization are used as the feature extraction module (Block2). The extraction module first uses a convolution kernel to extract features from the input feature matrix, and the extraction result is matrix-added to the input data and then normalized. The network processes the historical water level change vector and the historical precipitation change vector using the feature compression module with 512, 256, 128, and 64 convolution kernels respectively to obtain two 1*8*64 feature vectors. The key parameter vector is processed using the feature extraction module with 512, 256, 128, and 64 convolution kernels to obtain a 1*8*64 feature vector, and this feature vector is concatenated with the compressed feature vectors respectively to obtain two 1*8*128 feature vectors. Then, the two feature vectors are processed using the feature extraction module with 64 and 32 convolution kernels respectively to obtain two 1*8*32 feature vectors, and after concatenation, a 1*8*64 feature vector is obtained. Finally, the feature extraction module with 64 convolution kernels is used for high-order feature semantic extraction. After flattening, a fully connected layer with 512 is used for feature compression, and finally, a fully connected layer with a dimension of 1 is used to complete the final output.
[0088] The loss function of the reinforcement learning network designed by the present invention is shown in Figure 5, where and are the weight values of the flood discharge volume and the water storage volume in the total loss function respectively, and the range is 0-1; is the normalized predicted flood discharge volume and the normalized actual flood discharge volume; is the normalized predicted water storage volume and the normalized actual water storage volume; N is the total number of training space states in the training data; w1 is the weight parameter of the power generation reward, w2 is the weight parameter of the water supply reward, and w3 is the weight parameter of the flood control reward; is the normalized actual power generation volume and the normalized power generation volume demanded by the power grid; is the electricity price coefficient, and the range is 0-1; are the actual water supply volume and the minimum / maximum water supply demand respectively; is the current water level and the safety water level; is the flood control penalty coefficient, which is used to control the intensity of the flood control reward, and the range is 0-1.
[0089] The advantages of the present invention compared with the prior art are as follows:
[0090] (1) The present invention completes the rapid iteration of reinforcement learning by combining regression loss and state loss;
[0091] By combining regression loss and state loss, the present invention promotes the rapid iteration and optimization of the reinforcement learning model in reservoir operation. In the reservoir operation problem, there are certain conflicts and dependencies among multiple objectives (such as flood discharge, water storage, power generation, water supply, etc.). Traditional methods often have difficulty in balancing these objectives. However, the present invention effectively solves this problem through the combination of regression loss and state loss. The regression loss part focuses on the prediction accuracy of flood discharge and water storage. By minimizing the error between the predicted value and the true value, the model gradually improves the prediction accuracy of these two key parameters, making the scheduling decision more accurate. The state loss part combines the benefits of multiple objectives such as power generation, water supply, and flood control through the reward function, ensuring that while optimizing flood discharge and water storage, the model can also take into account the optimization of other objectives and avoid the over-priority of a certain objective. By dynamically adjusting the reward coefficient, the model can flexibly adapt to different scheduling requirements and find the optimal multi-objective balance. The design of combining regression loss and state loss makes the feedback in each training not only optimize the prediction accuracy but also improve the effect of multi-objective optimization, thus accelerating the rapid iteration process of reinforcement learning. Finally, the model can achieve intelligent scheduling decisions in continuous training to ensure the efficient utilization of reservoir resources.
[0092] (2) The present invention has high efficient dynamic reconstruction ability;
[0093] The present invention makes full use of data-driven modeling and decision-making methods to achieve precise and intelligent decision support in reservoir operation. Reservoir operation involves a large amount of dynamic data. Traditional methods often rely on expert experience or fixed rules for decision-making and are difficult to cope with complex environmental changes and diverse requirements. However, the present invention constructs a powerful data support system by combining real data with simulated data, enabling the model to be trained and optimized based on a large amount of historical data and automatically learning complex operation rules and decision-making patterns. In the data preprocessing stage, the present invention cleans and fills the missing values and outliers in the data to ensure that the model can effectively learn from the data during the training process. In addition, for features with different importance in the data, the invention assigns different weights to different data through hierarchical weighting to ensure that key factors can have a greater impact on the model during the training process, thereby improving the training efficiency and accuracy of the model. Combining these data-driven characteristics, the model can not only make adjustments based on real-time data but also predict future operation requirements according to the rules in historical data, realizing dynamic optimization and adaptive decision-making. This method makes reservoir operation more precise and efficient and can better handle various complex operation tasks in practical applications.
[0094] (3) The present invention has strong scalability and generality.
[0095] The present invention has strong scalability and generality and can be flexibly applied in various reservoir operation scenarios and can adapt to different types of reservoirs and resource operation problems. Reservoir operation problems are complex and diverse, and the challenges and requirements faced by different reservoirs are different. Traditional operation methods are often difficult to make flexible adjustments and optimizations. However, the present invention's deep learning-based reinforcement learning method can not only handle multi-objective optimization problems but also quickly adjust the structure and parameters of the model according to different requirements and environments, enabling the operation strategy to be adaptively optimized. By combining with reservoir operation simulation software, the model can be verified and tested in multiple different reservoir operation scenarios, thereby improving the generalization ability and practical application effect of the model. At the same time, the generality of this method is also manifested in its ability to handle different types of resource operation problems, such as power generation operation, water supply operation, flood management, etc. By simply adjusting the settings of the model input and loss function, the reinforcement learning model of the present invention can be applied to multiple fields to solve other similar resource operation problems. Therefore, this reinforcement learning-based reservoir operation method not only has high flexibility and can adapt to different tasks but also has strong expansion ability and can be effectively applied in multiple actual scenarios, providing an effective solution for broader resource management.
[0096] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A reservoir operation method based on multi-objective reinforcement learning, characterized in that Including: Step 1: Obtain the real data of information and reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir operation strategies, and external environment data. The reservoir operation strategies include flood discharge volume and water storage volume. Step 2: Based on the reservoir operation simulation software, build a reservoir operation model. Input the information into the reservoir operation model and output the simulated data of reservoir operation. Both the real data and the simulated data of reservoir operation include reservoir water level change, outflow discharge change, reservoir water storage change, power generation, water storage volume, flood water resource utilization rate, and water supply guarantee rate. Step 3: Preprocess the real data and the simulated data to obtain preprocessed data, and construct a data set according to the preprocessed data. The data set includes several groups of training space state data and corresponding reservoir operation strategies. The training space state data includes current water level, precipitation, inflow discharge, outflow discharge, temperature, historical water level change, historical precipitation change, power generation demand, and water use demand. Step 4: Input a group of the training space state data into a preset multi-objective deep reinforcement learning network and output the predicted reservoir operation strategy. Step 5: Input the reservoir structure characteristic parameters, reservoir historical operation data, and the predicted reservoir operation strategy into the reservoir operation model and output the predicted space state data. Step 6: Solve the regression distance between the predicted flood discharge volume and water storage volume and the real flood discharge volume and water storage volume to obtain the regression loss. Calculate the state loss according to the power generation, water supply volume, and current water level in the predicted space state data and the preset water supply demand and power generation demand through a reward function. Construct a loss function according to the regression loss and the state loss. Step 7: Update the parameters of the multi-objective deep reinforcement learning network based on the loss function in a backpropagation manner. Use the predicted space state data as the new training space state data and input it into the multi-objective deep reinforcement learning network with updated parameters. Iteratively update according to steps 5 to 7 until the regression loss is zero to obtain the updated target deep reinforcement learning network. Step 8: Input the training space state data of other groups in the several groups of training space state data into the target deep reinforcement learning network iteratively trained and updated according to steps 4 to 7 to obtain the trained multi-objective deep reinforcement learning network. Step 9: Input the real-time collected space state data into the trained multi-objective deep reinforcement learning network and output the real-time optimized reservoir operation strategy.
2. The reservoir operation method based on multi-objective reinforcement learning according to claim 1, wherein, Preprocessing the real data and the simulated data includes: Removing outliers, filling in missing values, and standardizing the real data and the simulated data, and classifying the data into three levels: general, important, and very important. The method for filling in missing values is the linear interpolation method, and the formula is: ; y is the interpolation result at the missing value position, is the abscissa of the known data points, is the ordinate of the known data points, and x is the abscissa of the missing value position; In the standardization process, the maximum-minimum normalization method is used for water level, inflow discharge, outflow discharge, temperature, flood discharge volume, and water storage volume, and the Z-score normalization method is used for precipitation, historical water level change, and historical precipitation change.
3. The reservoir operation method based on multi-objective reinforcement learning according to claim 1, wherein, The multi-objective deep reinforcement learning network has a three-input and two-output network structure. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with dimensions of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The multi-objective deep reinforcement learning network includes a feature compression module and a feature extraction module. The feature compression module is used to extract features from the input feature matrix using convolutional kernels, then superimpose the extraction results and the input data for matrix addition and perform normalization processing, and finally perform pooling processing. The feature extraction module is used to first extract features from the input feature matrix using convolutional kernels through the feature extraction module, and finally superimpose the extraction results and the input data for matrix addition and perform normalization processing. The multi-objective deep reinforcement learning network is used to process the historical water level change vector and the historical precipitation change vector respectively using the feature compression module with 512, 256, 128, and 64 convolutional kernels to obtain two compressed feature vectors of 1*8*64, and process the key parameter vector using the feature extraction module with 512, 256, 128, and 64 convolutional kernels to obtain a feature vector of 1*8*64. Concatenate this feature vector with the compressed feature vectors respectively to obtain two feature vectors of 1*8*128, then process the two feature vectors respectively using the feature extraction module with 64 and 32 convolutional kernels to obtain two feature vectors of 1*8*32, concatenate them to obtain a feature vector of 1*8*64, perform high-order feature semantic extraction using the feature extraction module with 64 convolutional kernels respectively, after flattening, use a fully connected layer with 512 for feature compression, and finally use a fully connected layer with a dimension of 1 to complete the final output.
4. The reservoir operation method based on multi-objective reinforcement learning according to claim 1, characterized in that, The loss function is as follows: ; Among them, and are the weight values of the control flood discharge and water storage in the total loss function, respectively, with a range of 0 - 1; are the normalized predicted flood discharge value and the normalized actual flood discharge value; are the normalized predicted water storage value and the normalized actual water storage value; N is the total number of training space state values in the training data; w1 is the weight parameter of the power generation reward, w2 is the weight parameter of the water supply reward, and w3 is the weight parameter of the flood control reward; are the normalized actual power generation value and the normalized power generation value demanded by the power grid; is the electricity price coefficient, with a range of 0 - 1; are the actual water supply, the minimum and maximum water supply demands, respectively; is the current water level and the safety water level; γ is the flood control penalty coefficient, used to control the intensity of the flood control reward, with a range of 0 - 1.
5. The reservoir scheduling device based on multi-objective reinforcement learning according to claim 1, characterized in that, It includes: An acquisition module, which is used to acquire information and the real data of reservoir operation. The information includes reservoir structure characteristic parameters, reservoir historical operation data, reservoir operation strategies, and external environment data. The reservoir operation strategies include flood discharge and water storage. A first output module, which is used to model a reservoir operation model based on reservoir operation simulation software, input the information into the reservoir operation model, and output the simulated data of reservoir operation. Both the real data and the simulated data of reservoir operation include reservoir water level change, outflow change, reservoir water storage change, power generation, water storage, flood water resource utilization rate, and water supply guarantee rate. A preprocessing module, which is used to preprocess the real data and the simulated data to obtain preprocessed data, and construct a data set according to the preprocessed data. The data set includes several groups of training space state data and corresponding reservoir operation strategies. The training space state data includes the current water level, precipitation, inflow, outflow, temperature, historical water level change, historical precipitation change, power generation demand, and water use demand. A second output module, which is configured to input a set of the training space state data into a preset multi-objective deep reinforcement learning network and output a predicted reservoir operation strategy; A third output module, which is configured to input the reservoir structure feature parameters, the reservoir historical operation data, and the predicted reservoir operation strategy into the reservoir operation model and output predicted space state data; A construction module, which is configured to solve a regression distance between the predicted flood discharge and water storage and the actual flood discharge and water storage to obtain a regression loss, calculate a state loss through a reward function based on the generated electricity, water supply, and current water level in the predicted space state data and preset water supply demands and power generation demands, and construct a loss function according to the regression loss and the state loss; An update module, which is configured to update the parameters of the multi-objective deep reinforcement learning network based on the loss function in a backpropagation manner, use the predicted space state data as new training space state data, input the data into the multi-objective deep reinforcement learning network with updated parameters, and iterate and update according to steps five to seven until the regression loss is zero to obtain an updated target deep reinforcement learning network; A training module, which is configured to input the training space state data of other groups in the several groups of training space state data into the target deep reinforcement learning network that is iteratively trained and updated according to steps four to eight to obtain a trained multi-objective deep reinforcement learning network; An optimization strategy module, which is configured to input the real-time collected space state data into the trained multi-objective deep reinforcement learning network and output a real-time optimized reservoir operation strategy.
6. The reservoir scheduling device based on multi-objective reinforcement learning according to claim 5, characterized in that, Preprocess the real data and the simulation data, including: Removing outliers, filling in missing values, and performing normalization processing on the real data and the simulation data, and classifying the data into three levels: general, important, and very important; The method for filling in missing values is the linear interpolation method, and the formula is: ; y is the interpolation result at the missing value position, is the abscissa of the known data points, is the ordinate of the known data points, and x is the abscissa of the missing value position; In the normalization processing, the maximum-minimum normalization method is used for the water level, the inflow, the outflow, the temperature, the flood discharge, and the water storage, and the Z-score normalization method is used for the precipitation, the historical water level change, and the historical precipitation change.
7. The reservoir scheduling device based on multi-objective reinforcement learning according to claim 5, characterized in that, The multi-objective deep reinforcement learning network has a network structure with three inputs and two outputs. Among them, the three inputs are the historical water level change vector, the historical precipitation change vector, and the key parameter vector, with dimensions of 1*128*1, 1*128*1, and 1*8*1 respectively. The key parameter vector includes the current water level, precipitation, inflow, outflow, temperature, humidity, power generation demand, and water use demand. The multi-objective deep reinforcement learning network includes a feature compression module and a feature extraction module. The feature compression module is used to extract features from the input feature matrix using a convolutional kernel, then stack the extraction result and the input data for matrix addition and perform normalization processing, and finally perform pooling processing. The feature extraction module is used to first extract features from the input feature matrix using a convolutional kernel through the feature extraction module, and finally stack the extraction result and the input data for matrix addition and perform normalization processing. The multi-objective deep reinforcement learning network is used to process the historical water level change vector and the historical precipitation change vector respectively using the feature compression module with 512, 256, 128, and 64 convolutional kernels to obtain two compressed feature vectors of 1*8*64, process the key parameter vector using the feature extraction module with 512, 256, 128, and 64 convolutional kernels to obtain a feature vector of 1*8*64, splice the feature vector with the compressed feature vectors respectively to obtain two feature vectors of 1*8*128, then process the two feature vectors respectively using the feature extraction module with 64 and 32 convolutional kernels to obtain two feature vectors of 1*8*32, splice them to obtain a feature vector of 1*8*64, perform high-order feature semantic extraction using the feature extraction module with 64 convolutional kernels respectively, after flattening, use a fully connected layer with 512 for feature compression, and finally use a fully connected layer with a dimension of 1 to complete the final output.
8. A reservoir scheduling device based on multi-objective reinforcement learning according to claim 5, characterized in that The loss function is as follows: ; Among them, and are the weight values of the control flood discharge and water storage in the total loss function, respectively, with the range of 0 - 1; are the normalized predicted flood discharge value and the normalized actual flood discharge value; are the normalized predicted water storage value and the normalized actual water storage value; N is the total number of training space state values in the training data; w1 is the weight parameter of power generation reward, w2 is the weight parameter of water supply reward, and w3 is the weight parameter of flood control reward; are the normalized actual power generation value and the normalized power generation value demanded by the power grid; is the electricity price coefficient, with the range of 0 - 1; are the actual water supply, the minimum and maximum water supply demands, respectively; are the current water level and the safety water level; is the flood control penalty coefficient, used to control the intensity of flood control reward, with the range of 0 - 1, represents the water supply guarantee rate, represents the utilization rate of flood resources.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
Reinforcement learning model FQI-based reservoir flood control optimal scheduling method
CN112966445A
Reservoir group scheduling decision behavior mining method and reservoir scheduling automatic control device
CN113204583A
Reservoir level prediction and early warning method based on neural network and GCN deep learning model
CN115310536A
Disposal method for events influencing cascade reservoir dispatching operation
CN119514922A
Cited By
Adaptive reservoir scheduling method and device based on deep learning, equipment and medium
CN121146955A