Method for predicting tail water level of hydraulic power plant based on reinforcement learning
By combining upstream and downstream operation data and environmental data of hydropower plants, the tailwater level prediction model is optimized using Attention-LSTM and DQN reinforcement learning algorithms, the problem of low tailwater level prediction accuracy in the existing technology is solved, and real-time and accurate tailwater level prediction is achieved.
Patent Information
- Application Number
- CN202510877844.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to make full use of various related data in the operation of hydropower plants, resulting in low accuracy of water level prediction of hydropower plants, affecting the economic benefits and safe operation of power plants.
Using a reinforcement learning-based method, combining upstream and downstream operation data and environmental data of hydropower plants, the tailwater level prediction model is optimized through the Attention-LSTM model and the DQN reinforcement learning algorithm, state space, action space and reward functions are constructed, and model parameters are adaptively adjusted to improve prediction accuracy.
It enhances the robustness and accuracy of tailwater level prediction, can adapt to complex operating environments, and provides real-time and accurate tailwater level prediction results.
Smart Images

Figure CN120387556A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tail water level prediction of hydropower plants, and particularly relates to a method for predicting the tail water level of hydropower plants based on reinforcement learning. Background Technique
[0002] The tail water level of a hydropower station refers to the river water level at the outlet of the draft tube of the hydropower station powerhouse, and it is one of the important parameters for determining the working head of the hydropower station. The methods for predicting the tail water level of hydropower plants often have difficulty in making full use of various associated data available during the operation of hydropower plants, which results in low prediction accuracy and affects the economic benefits and safe operation of the power plants.
[0003] In the prior art, CN115222165B proposes a method and system for predicting the operating state of a drainage system based on a Transformer model. This patent realizes water level prediction by using historical operating parameters and future transportation parameters for model prediction. However, in actual situations, the factors affecting water level prediction are numerous and complex, and simply considering operating parameters cannot accurately predict future water level conditions.
[0004] CN113344288B proposes a method, device and computer-readable storage medium for predicting the water level of cascade hydropower stations. This patent can use the collected hydrological information and meteorological information to assist in predicting the water level information of hydropower stations, but this technology cannot make full use of the operation data of hydropower stations, resulting in the inability to solve the water level prediction task in a complex operating environment.
[0005] The current technical difficulty lies in that, in the case of being able to obtain various associated data during the operation of hydropower plants, the associated data between upstream and downstream cannot be fully processed, and complex environmental changes cannot be dealt with, which affects the tail water level prediction work of hydropower plants and seriously affects the tail water level prediction results. Summary of the Invention
[0006] The present invention provides a method for predicting the tail water level of hydropower plants based on reinforcement learning to solve the problems raised in the above background technique.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is: A method for predicting the tail water level of hydropower plants based on reinforcement learning, comprising the following steps: Step 1: Obtain the current data and historical data of the hydropower plant; the current data includes the current upstream and downstream operation data of the hydropower plant, the current environmental data, and the current power generation plan; the historical data includes: historical power generation plans, historical hydropower plant environmental data, historical hydropower plant upstream and downstream operation data, and historical hydropower plant tail water level data; Step 2: Build a tail water level impact prediction model, train the tail water level impact prediction model using historical data, and input the current data into the trained model to obtain environmental impact factors and operation impact factors; Step 3: Build a tail water level prediction model based on Attention-LSTM, train the tail water level prediction model using historical data to obtain a pre-trained tail water level prediction model; Step 4: Model the hydropower plant operation environment as a Markov decision process, define the state space, action space, and reward function in combination with environmental impact factors and operation impact factors; use the DQN reinforcement learning algorithm based on the attention mechanism as the core of the intelligent agent to optimize the parameters of the pre-trained tail water level prediction model to obtain an optimized tail water level prediction model; Step 5: Use the adjusted Attention-LSTM model to predict the tail water level to obtain the tail water level prediction result. Furthermore, in Step 2, the training process of the tail water level impact prediction model includes: Initialize a weak learner; Input the training data into the model for iterative training. During this period, calculate the residual value between the true value and the predicted value of the training sample, and train a new decision tree with the residual value as the target variable; Add the newly trained decision tree to the model and update the predicted value of the model; When the maximum number of iterations is reached or the performance of the validation set no longer improves, stop the iterative training and evaluate the model performance using the mean square error; Before iterative training, use the grid search algorithm to select the optimal parameters of the gradient boosting tree, and perform iterative training in combination with the optimal parameters.
[0008] Furthermore, in Step 3, the structure of the tail water level prediction model based on Attention-LSTM includes: Data input layer, used to receive preprocessed historical hydropower plant environment data and historical hydropower plant operation data; LSTM feature extraction layer, used to extract the temporal features of time series data through the LSTM network; Attention weighting layer, used to calculate the importance weights of each time step and weight the output of the LSTM feature extraction layer to highlight the influence of key time steps; Result output layer, used to integrate the weighted LSTM output and output the final tail water level prediction result.
[0009] Furthermore, the LSTM feature extraction layer specifically includes the following operation steps: The LSTM feature extraction layer receives the data transmitted by the data input layer step by step in time; Each LSTM cell updates its internal state based on the current input, the hidden state and the memory state at the previous moment, and outputs the current hidden state; By stacking multiple LSTM cells, the tail water level prediction model captures the temporal dependence relationships at different time scales; The LSTM feature extraction layer outputs the hidden state sequence at each time step to the attention weighting layer.
[0010] Furthermore, the attention weighting layer specifically includes the following operation steps: The attention weighting layer receives the hidden state sequence output by the LSTM feature extraction layer; For each time step, the attention weighting layer calculates an attention weight; Multiply each hidden state by its corresponding attention weight to obtain the weighted hidden state; The attention weighting layer outputs the weighted hidden state to the result output layer.
[0011] Furthermore, in the step of "For each time step, the attention weighting layer calculates an attention weight", the specific process includes the following: using the Transformer network to map each hidden state into an attention score, and then using the softmax function to normalize the attention score into a weight.
[0012] Furthermore, in step four, the state space includes: the upstream and downstream operation data of the hydropower plant , environmental data , historical tail water level data , tail water level prediction error , environmental impact factor and operation impact factor ; The action space includes a series of discrete actions of model parameters.
[0013] Furthermore, in step four, the reward function is specifically expressed as: ; where , and are weight coefficients, is the tail water level prediction error, is the environmental impact factor, is the operation impact factor.
[0014] Furthermore, in step four, the core components of the DQN algorithm based on the attention mechanism include: The attention Q network, which adds an attention mechanism to the original Q network, and the Q value of executing a certain action under a given state, that is, the expected cumulative reward; A target attention Q network for processing time series data in the input state; An experience replay buffer for storing the interaction experiences between the agent and the environment.
[0015] Furthermore, in step four, the specific training process of the DQN reinforcement learning algorithm based on the attention mechanism includes: Initialize the attention Q network, the target attention Q network, and the experience replay buffer; The agent selects an action from the action space according to the current state variable and the attention Q network, and executes it, adjusting the parameters of the Attention-LSTM model; Obtain the reward according to the tail water level prediction error, the environmental impact factor, and the operation impact factor calculated by the adjusted model; The agent observes the next state after executing the action according to the reward, and stores the obtained experience in the experience replay buffer; Randomly sample a batch of experiences from the experience replay buffer for training the attention Q network. During the training, use the target attention Q network to calculate the maximum Q value of the next state, and use the gradient descent algorithm to minimize the error between the predicted Q value and the target Q value of the attention Q network until the attention Q network converges, that is, the change of the Q value tends to be stable.
[0016] The present invention can achieve the following beneficial effects: 1. The present invention combines the upstream and downstream operation data and environmental data for the tail water level prediction of the hydropower plant, which not only enhances the dimension of the data features, but also considers the influence of the correlation of the upstream and downstream operation data on the tail water level prediction, enhances the robustness of the tail water level prediction, provides a data basis for the subsequent tail water level impact prediction and tail water level prediction, and improves the accuracy of the tail water level prediction.
[0017] 2. The present invention uses the tail water level impact prediction model to learn the influence degree of the upstream and downstream operation data and environmental data on the tail water level prediction of the hydropower plant, and can dynamically adjust the tail water level prediction model by using the obtained real-time operation impact factor and environmental impact factor, so that the model can adapt to the complex operation environment and improve the accuracy of the tail water level prediction.
[0018] 3. The present invention combines the tail water level prediction model based on Attention-LSTM with the DQN reinforcement learning algorithm based on the attention mechanism, optimizes the parameters of the tail water level prediction model by defining the state space, action space, and reward function of the Markov decision-making, obtains the tail water level prediction model that conforms to the specific situation under the actual operation environment, and obtains the tail water level prediction result according to the optimized tail water level prediction model, thereby improving the accuracy of the tail water level prediction. Description of the Drawings
[0019] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 It is a flowchart of a method for predicting the tail water level of a hydropower plant based on reinforcement learning according to the present invention; Figure 2 It is the structure of the tail water level prediction model of the present invention. Specific embodiments
[0020] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant accompanying drawings. Embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0021] As Figure 1 and Figure 2 shown, a method for predicting the tail water level of a hydropower plant based on reinforcement learning includes the following steps: Step 1: Obtain the current data and historical data of the hydropower plant from the background system of the hydropower plant. The current data is used to obtain the tail water level prediction result, including the current upstream and downstream operation data of the hydropower plant, the current environmental data, and the current power generation plan. The historical data is used for the training of the tail water level influence prediction model and the tail water level prediction model, including the historical power generation plan, the historical hydropower plant environmental data, the historical hydropower plant upstream and downstream operation data, and the historical hydropower plant tail water level data.
[0022] Among them, the current power generation plan and the historical power generation plan belong to the hydropower plant management data; the hydropower plant upstream and downstream operation data includes: unit output power, gate opening, diversion flow, equipment status data, etc.; the environmental data includes: meteorological data and hydrological data. The meteorological data includes: rainfall, temperature, humidity, wind speed, wind direction, etc.
[0023] Step 2: Construct a tail water level influence prediction model. The tail water level influence prediction model is used to obtain the influence degree of environmental data and upstream and downstream operation data on the tail water level prediction. The tail water level influence prediction model is trained by using the historical power generation plan, the historical hydropower plant environmental data, the historical hydropower plant upstream and downstream operation data, and the historical hydropower plant tail water level data to obtain a trained tail water level influence prediction model; then, the current power generation plan, the current hydropower plant operation data, and the current environmental data are input into the trained model to obtain an environmental influence factor and an operation influence factor .
[0024] The tail water level impact prediction model selects the gradient boosting tree model, which constructs a strong prediction model by iteratively training a series of decision trees. The specific training process of this model is as follows: Initialize a weak learner; input the training data into the model for iterative training. During this process, calculate the residual value between the true value and the predicted value of the training samples, and use the residual value as the target variable to train a new decision tree to fit the residuals as much as possible; add the newly trained decision tree to the model and update the predicted value of the model; when the maximum number of iterations is reached or the performance of the validation set no longer improves, stop the iterative training and evaluate the model performance using the mean squared error; in order to improve the convergence speed of the tail water level impact prediction model, use the grid search algorithm to select the optimal parameters of the gradient boosting tree before iterative training, and perform iterative training in combination with the optimal parameters.
[0025] In addition, the environmental impact factor and the operation impact factor are both between [0, 1]. The larger the value, the greater the predicted impact of the environmental factor or the operation factor on the tail water level. The environmental impact factor and the operation impact factor are used for subsequent parameter optimization and adjustment of the tail water level prediction model of the hydropower plant; Step 3: Construct a tail water level prediction model based on Attention-LSTM. The tail water level prediction model is used to obtain the tail water level prediction result based on the upstream and downstream operation data and environmental data. By inputting the historical hydropower plant environmental data, historical hydropower plant upstream and downstream operation data, and historical hydropower plant tail water level data into the tail water level prediction model based on Attention-LSTM for training, a pre-trained tail water level prediction model is obtained, which provides a basic model for subsequent parameter optimization of the tail water level prediction model. The pre-trained tail water level prediction model takes the parameters in the state space as input and outputs the tail water level prediction data.
[0026] The structure of the tail water level prediction model based on Attention-LSTM includes: a data input layer, an LSTM feature extraction layer, an attention weighting layer, and a result output layer.
[0027] Among them, the data input layer is used to receive the preprocessed historical hydropower plant environmental data and historical hydropower plant operation data.
[0028] The LSTM feature extraction layer is used to extract the temporal features of time series data through the LSTM network; the LSTM feature extraction layer receives the data passed from the data input layer step by step in time; each LSTM unit updates its internal state according to the current input, the hidden state of the previous moment, and the memory state, and outputs the current hidden state; through the stacking of multiple LSTM units, the tail water level prediction model can capture the temporal dependencies at different time scales; the LSTM feature extraction layer finally outputs the hidden state sequence of each time step to the attention weighting layer.
[0029] The attention weighting layer is used to calculate the importance weights for each time step and weight the output of the LSTM feature extraction layer to highlight the influence of key time steps. The attention weighting layer receives the sequence of hidden states output by the LSTM feature extraction layer; for each time step, the attention weighting layer calculates an attention weight indicating the contribution degree of that time step to the prediction result. The process is as follows: Use the Transformer network to map each hidden state to an attention score, and then use the softmax function to normalize the attention scores into weights. Subsequently, multiply each hidden state by its corresponding attention weight to obtain the weighted hidden state; the attention weighting layer outputs the weighted hidden state to the result output layer.
[0030] The result output layer is used to integrate the weighted LSTM output and output the final tail water level prediction result; the result output layer receives the weighted hidden state output by the attention weighting layer, maps the weighted hidden state to the tail water level prediction value through the use of a fully connected layer, and then outputs it.
[0031] Since the parameters of the tail water level prediction model based on Attention-LSTM are fixedly set during the initial training and not optimized in combination with the actual situation, for the complex operating environment of the hydropower plant, it is necessary to adaptively adjust the parameters of the tail water level prediction model; Step 4: In order to optimize the parameters of the tail water level prediction model to improve the prediction accuracy, model the operating environment of the hydropower plant as a Markov decision process. The state space consists of the upstream and downstream operating data and environmental data of the hydropower plant. The action space is defined as the parameter adjustment of the tail water level prediction model, and the reward function is designed as a function based on the prediction error, environmental impact factor, and operating impact factor; use the DQN reinforcement learning algorithm based on the attention mechanism as the core of the intelligent agent to optimize the parameters of the pre-trained tail water level prediction model.
[0032] In addition to the upstream and downstream operating data of the hydropower plant in the state space of the Markov decision and environmental data there are also historical tail water level data , tail water level prediction error , environmental impact factor and operating impact factor ; among them, the tail water level prediction error , environmental impact factor and operating impact factor also constitute the reward function of the Markov decision process.
[0033] The action space includes a series of discrete actions of model parameters, such as: increasing the learning rate, decreasing the learning rate, increasing the number of LSTM layers, decreasing the number of LSTM layers, increasing the batch size, decreasing the batch size, remaining unchanged, etc.
[0034] The reward function can be expressed as: ; , and are weight coefficients used to balance the tail water level prediction error, environmental impact factor, and operation impact factor. The design of the reward function is to encourage the tail water level prediction model to reduce the prediction error and adjust the model parameters according to the environmental impact factor and operation impact factor. Among them, the environmental impact factor and operation impact factor are obtained by inputting the current upstream and downstream operation data and current environmental data into the trained tail water level impact prediction model.
[0035] Taking the DQN reinforcement learning algorithm based on the attention mechanism as the agent core is to learn the optimal adjustment strategy from the complex operation environment of the hydropower plant, so as to adaptively adjust the parameters of the Attention-LSTM tail water level prediction model to achieve higher tail water level prediction accuracy.
[0036] The core components of the DQN algorithm based on the attention mechanism include: attention Q network, target attention Q network, and experience replay buffer. The attention Q network adds an attention mechanism to the original Q network, and the Q value of performing a certain action in a given state, that is, the expected cumulative reward; the attention mechanism is used to process the time series data in the input state to better capture the time-dependent relationship in the state. The target attention Q network is a copy of the attention Q network, which is used to calculate the target Q value to improve the stability of model training. The experience replay buffer is used to store the interaction experience between the agent and the environment; the experience includes state, action, reward, and next state; experience replay can provide the agent with sampled experience for training the attention Q network.
[0037] The specific training process of the DQN reinforcement learning algorithm based on the attention mechanism includes: initializing the attention Q network, target attention Q network, and experience replay buffer; the agent selects an action from the action space using the ε-greedy strategy according to the current state variable and attention Q network and executes it to adjust the parameters of the Attention-LSTM model; obtaining the reward according to the tail water level prediction error, environmental impact factor, and operation impact factor calculated by the adjusted model; the agent observes the next state after executing the action according to the reward and stores the obtained experience in the experience replay buffer; randomly sampling a batch of experiences from the experience replay buffer for training the attention Q network, calculating the maximum Q value of the next state using the target attention Q network during training, and using the gradient descent algorithm to minimize the error between the predicted Q value and the target Q value of the attention Q network; until the attention Q network converges, that is, the change of the Q value tends to be stable.
[0038] Step 5: Use the adjusted Attention-LSTM model to predict the tail water level and obtain the tail water level prediction result.
[0039] In summary, this application provides an intelligent prediction method that can make full use of the operation data of hydropower plants, predict the tail water level in real time and accurately, and can adapt to the complex operation environment of hydropower plants. The prediction method of this application has the following advantages: Predict the influence degree of the tail water level by combining the upstream and downstream operation data, environmental data and management data of the hydropower plant. By analyzing the upstream and downstream operation data and environmental data of the hydropower plant, and using the model to learn the correlation between the operation data upstream and downstream and the correlation between the operation environment data and the power generation plan, it makes up for the single problem of the existing technology and ensures the comprehensiveness of the influencing factors of the tail water level; these data provide reference data for the subsequent tail water level influence prediction model and tail water level prediction model.
[0040] Through the tail water level influence prediction model, not only can the influence degree of the environment on the predicted tail water level be learned, but also the influence brought by the operation data of upstream and downstream hydropower stations can be predicted; at the same time, the real-time influence factor generated by the tail water level influence prediction model can be used to adjust the parameters of the tail water level prediction model in real time to cope with the complex operation environment, thereby improving the accuracy of the tail water level prediction.
[0041] Use the tail water level prediction model based on Attention-LSTM to take the upstream and downstream operation data and environmental data of the hydropower plant as input and output the tail water level prediction result; the structure of this model includes: data input layer, LSTM feature extraction layer, attention weighting layer and result output layer; use the LSTM feature extraction layer to extract the long-distance dependence relationship between data, and use the attention weighting layer to assign weights to the importance of data features in each time step; combining multi-source data and model structure can effectively improve the prediction accuracy of the tail water level of the hydropower plant.
[0042] In order to overcome the influence of the complex and changeable operation environment on the tail water level prediction model with fixed parameters, model the operation environment of the hydropower plant as a Markov decision process, define the state space, action space and reward function, and use the DQN reinforcement learning algorithm based on the attention mechanism as the core of the intelligent agent to optimize the parameters of the tail water level prediction model, so as to ensure that the tail water level prediction model can output accurate tail water level prediction results.
[0043] When facing the changing working environment of the hydropower plant, the collected operation data such as the power generation plan and gate opening of the hydropower plant and the environmental data are input into the trained tail water level influence prediction model, tail water level prediction model and the DQN network based on the attention mechanism. The tail water level prediction model adjusted by the Markov decision process is used to predict the tail water level, so as to obtain the tail water level prediction result that conforms to the real-time operation environment, thus ensuring the accuracy and reliability of the tail water level prediction of the hydropower plant.
[0044] It should be noted that the abbreviations in the above text are all existing expressions. For the convenience of understanding, they are further described here. Among them, Attention-LSTM is the attention long short-term memory network, LSTM is the long short-term memory network, and DQN is the deep Q network.
[0045] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for predicting the tail water level of a hydropower plant based on reinforcement learning, characterized in that It includes the following steps: Step 1: Obtain the current data and historical data of the hydropower plant; the current data includes the upstream and downstream operation data of the current hydropower plant, the current environmental data, and the current power generation plan; the historical data includes: historical power generation plans, historical hydropower plant environmental data, historical upstream and downstream operation data of the hydropower plant, and historical tail water level data of the hydropower plant; Step 2: Construct a tail water level impact prediction model, train the tail water level impact prediction model using historical data, and input the current data into the trained model to obtain environmental impact factors and operation impact factors; Step 3: Construct a tail water level prediction model based on Attention-LSTM, train the tail water level prediction model using historical data to obtain a pre-trained tail water level prediction model; Step 4: Model the operation environment of the hydropower plant as a Markov decision process, define the state space, action space, and reward function in combination with the environmental impact factors and operation impact factors; use the DQN reinforcement learning algorithm based on the attention mechanism as the core of the intelligent agent to optimize the parameters of the pre-trained tail water level prediction model to obtain an optimized tail water level prediction model; Step 5: Use the adjusted Attention-LSTM model to predict the tail water level to obtain the tail water level prediction result.
2. The method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 1, wherein: In Step 2, the training process of the tail water level impact prediction model includes: Initialize a weak learner; Input the training data into the model for iterative training. During this period, calculate the residual value between the true value and the predicted value of the training samples, and train a new decision tree with the residual value as the target variable; Add the newly trained decision tree to the model and update the predicted value of the model; When the maximum number of iterations is reached or the performance of the validation set no longer improves, stop the iterative training and evaluate the model performance using the mean square error; Before iterative training, use the grid search algorithm to select the optimal parameters of the gradient boosting tree and perform iterative training in combination with the optimal parameters.
3. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 1, characterized in that: In Step 3, the structure of the tail water level prediction model based on Attention-LSTM includes: Data input layer, used to receive the pre-processed historical hydropower plant environmental data and historical hydropower plant operation data; LSTM feature extraction layer, used to extract the temporal features of time series data through the LSTM network; Attention weighting layer, used to calculate the importance weights of each time step and weight the output of the LSTM feature extraction layer to highlight the influence of key time steps; Result output layer, used to integrate the weighted LSTM output and output the final tail water level prediction result.
4. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 3, characterized in that: The LSTM feature extraction layer specifically includes the following operation steps: The LSTM feature extraction layer receives the data transmitted by the data input layer step by step in time; Each LSTM unit updates its internal state according to the current input, the hidden state and memory state of the previous moment, and outputs the current hidden state; Through the stacking of multiple LSTM units, the tail water level prediction model captures the temporal dependence relationships at different time scales; The LSTM feature extraction layer outputs the hidden state sequence of each time step to the attention weighting layer.
5. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 4, characterized in that: The attention weighting layer specifically includes the following operation steps: The attention weighting layer receives the hidden state sequence output by the LSTM feature extraction layer; For each time step, the attention-weighted layer calculates an attention weight; Multiply each hidden state by its corresponding attention weight to obtain the weighted hidden state; The attention-weighted layer outputs the weighted hidden state to the result output layer.
6. The method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 5, wherein: In the step of "For each time step, the attention-weighted layer calculates an attention weight", specifically it includes the following process: using the Transformer network to map each hidden state to an attention score, and then using the softmax function to normalize the attention score into a weight.
7. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 3, characterized in that: In step four, the state space includes: operation data of the upstream and downstream of the hydropower plant , environmental data , historical tail water level data , tail water level prediction error , environmental impact factor , and operation impact factor ; The action space includes a series of discrete actions of model parameters.
8. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 7, characterized in that: In step four, the reward function is specifically expressed as: ; Among them, , and are weight coefficients, is the prediction error of the tail water level, is the environmental impact factor, is the operation impact factor.
9. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 1, characterized in that: In step four, the core components of the DQN algorithm based on the attention mechanism include: The attention Q-network, which adds an attention mechanism to the original Q-network, and the Q value of performing a certain action under a given state, that is, the expected cumulative reward; The target attention Q-network, which is used to process the time series data in the input state; The experience replay buffer, which is used to store the interaction experience between the agent and the environment.
10. A method for predicting the tail water level of a hydropower plant based on reinforcement learning according to claim 9, characterized in that: In step four, the specific training process of the DQN reinforcement learning algorithm based on the attention mechanism includes: Initialize the attention Q-network, the target attention Q-network, and the experience replay buffer; The agent selects an action from the action space using the ε-greedy strategy according to the current state variable and the attention Q-network and executes it, adjusting the parameters of the Attention-LSTM model; Obtain the reward according to the tail water level prediction error, environmental impact factor, and operation impact factor calculated by the adjusted model; The agent observes the next state after performing the action according to the reward and stores the obtained experience in the experience replay buffer; Randomly sample a batch of experiences from the experience replay buffer for training the attention Q-network. During training, use the target attention Q-network to calculate the maximum Q value of the next state, and use the gradient descent algorithm to minimize the error between the predicted Q value and the target Q value of the attention Q-network until the attention Q-network converges, that is, the change of the Q value tends to be stable.
Citation Information
Patent Citations
Hydropower station in-plant real-time optimization scheduling method based on depth enhancement algorithm
CN114841595A
Water level prediction method based on whale optimization algorithm and long-short term memory network
CN114912673A
Pumped storage power station operation trend early warning analysis method and system
CN119106239A
Real-time monitoring system and method of plasma generator
CN119739970A