A pumped storage power station carbon optimization method based on deep reinforcement learning

By using a multi-agent system based on deep reinforcement learning and a dynamic reward function, the problems of accuracy and real-time performance in energy equipment health monitoring have been solved, enabling efficient and intelligent operation of pumped storage power stations and improving the reliability and economy of equipment operation.

CN120875257BActive Publication Date: 2026-01-27POWERCHINA BEIJING ENG CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510994957.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-01-27
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing methods for monitoring the health of energy equipment are insufficient in terms of monitoring accuracy, real-time performance, and intelligence, making it difficult to meet the monitoring needs under complex operating conditions.

Method used

A multi-agent system based on deep reinforcement learning is adopted. By fusing multi-source data, designing dynamic reward functions, and real-time monitoring, combined with deep reinforcement learning algorithms and physical models, carbon emission optimization and real-time adjustment of equipment status of pumped storage power stations can be achieved.

Benefits of technology

It significantly improves the accuracy and efficiency of health monitoring for energy equipment, enabling rapid response to changes in equipment status, early identification of complex failure modes, extension of equipment lifespan, and reduction of maintenance costs, ensuring efficient operation of the system under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875257B_ABST
    Figure CN120875257B_ABST
Patent Text Reader

Abstract

The application provides a pumped storage power station carbon optimization method based on deep reinforcement learning, relates to the technical field of energy equipment management and optimization, and comprises the following steps: collecting and preprocessing data; calling a multi-agent system; designing a dynamic reward function; adopting a deep reinforcement learning algorithm to train a strategy prediction model of each agent in the multi-agent system by taking the best carbon emission effect in the latest period T as an optimization target; and applying an optimization strategy and adjusting in real time to ensure stable operation of the system. Through the synergistic effect of each step, the application realizes low-carbon and efficient operation of the pumped storage power station, and has significant economic and environmental benefits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy equipment management and optimization technology, specifically to a carbon optimization method for pumped storage power plants based on deep reinforcement learning. Background Technology

[0002] In recent years, with the rapid development of the energy industry, health monitoring technology for energy equipment has gradually become a research hotspot. Traditional monitoring methods mainly rely on manual inspections and periodic maintenance, which are inefficient and make it difficult to detect potential faults in a timely manner. With the advancement of sensor technology and the rise of big data analytics, data-driven health monitoring methods are gradually emerging. These methods significantly improve the efficiency and accuracy of monitoring by collecting equipment operating data and using statistical analysis and machine learning algorithms for fault diagnosis and prediction. However, existing technologies still have many shortcomings in terms of real-time performance, accuracy, and intelligence.

[0003] Despite the progress made in data-driven health monitoring technologies, existing methods still suffer from the following problems: First, traditional methods often rely on a single data source and lack the ability to fuse and process multi-source heterogeneous data, resulting in incomplete and inaccurate monitoring results. Second, existing technologies are insufficient in terms of real-time performance, failing to respond quickly to dynamic changes in equipment status and struggling to meet monitoring needs under complex operating conditions. Furthermore, the level of intelligence in existing methods needs improvement, particularly in the identification and prediction of complex fault modes, where adaptive learning capabilities are lacking. These issues limit the effectiveness of existing technologies in practical applications, making it difficult to achieve efficient health monitoring of energy equipment. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a carbon optimization method for pumped storage power stations based on deep reinforcement learning, which solves the problems of insufficient monitoring accuracy, poor real-time performance, and lack of intelligent dynamic monitoring in existing energy equipment health monitoring methods.

[0005] The technical solution adopted in this invention is as follows:

[0006] This invention provides a carbon optimization method for pumped storage power plants based on deep reinforcement learning, comprising the following steps:

[0007] Step S1: Collect data on the pumped storage power station at the current time. The real-time operating data is collected and preprocessed to obtain preprocessed real-time operating data; wherein, the preprocessed real-time operating data includes the pumped storage power station at the current time. Water level, flow rate, equipment status and carbon emissions ;

[0008] Step S2: Invoke the multi-agent system; the multi-agent system is a cooperative system composed of multiple agents; each agent is assigned a specific optimization task; agents share state space and action space through a predefined communication protocol; each agent has a policy prediction model, and each agent generates the final policy through the policy prediction model.

[0009] Step S3, Design the dynamic reward function The dynamic reward function With the current moment carbon emissions It is related to the operating efficiency of pumped storage power stations;

[0010] Step S4: Using a deep reinforcement learning algorithm, with the optimization objective of achieving the best carbon emission performance based on the most recent period T, the policy prediction model of each agent in the multi-agent system is trained. During training, the agents select actions in the action space based on their current state and according to the dynamic reward function. Update the policy prediction model, and then make a decision on the current time based on the updated policy prediction model. Strategies;

[0011] Step S5, for each agent at the current time A comprehensive analysis of the strategies is conducted to determine the current moment. A comprehensive strategy;

[0012] Step S6, based on the current time A comprehensive strategy to dynamically adjust the pumped storage power station at the current moment. The operating parameters are adjusted to minimize carbon emissions, and the current time after the operating parameters are adjusted is monitored in real time. carbon emissions The adjustment factor is obtained using the following formula. :

[0013]

[0014] in: For adjustment coefficients;

[0015] According to the adjustment factor The policy prediction model for each agent is dynamically adjusted to obtain the dynamically adjusted policy prediction model.

[0016] Step S7, let Return to step S1.

[0017] Preferably, the real-time operating data also includes grid load demand, renewable energy generation, and carbon trading market prices.

[0018] Preferably, in step S1, the real-time running data is preprocessed, including:

[0019] The real-time operating data is subjected to noise filtering to remove abnormal values ​​caused by sensor errors or environmental interference.

[0020] The real-time running data is preprocessed using a normalization function to obtain preprocessed real-time running data. The normalization method is as follows:

[0021] Read the running data of the most recent period T, and determine the minimum value among the running data of the most recent period T. and maximum value ; the current moment Real-time operational data is represented as The following formula is used to... Normalization is performed to obtain normalized real-time running data. ;

[0022]

[0023] The preprocessed real-time running data is stored in a temporary buffer.

[0024] Preferably, in step S2, the state space is defined as follows: Action space is defined as ;in, For the state space a state; In the action space A type of action.

[0025] Preferably, the multi-agent system further includes a fault prediction module for predicting potential equipment faults in advance; the fault prediction module uses a deep learning algorithm to make predictions based on historical fault data and real-time operating data.

[0026] Preferably, the dynamic reward function is:

[0027]

[0028] in:

[0029] , and The weighting coefficients are calculated using historical data normalization and optimized using a genetic algorithm.

[0030] For a moment The equipment efficiency of a pumped storage power station is given by the formula: ; )and , respectively time The device's output power and input power;

[0031] For a moment The economic indicators of pumped storage power stations are given by the following formula: ; and , respectively time The revenue and costs of pumped storage power stations;

[0032] For a moment The carbon emissions of pumped storage power stations are collected in real time through sensors;

[0033] or

[0034] The dynamic reward function is:

[0035]

[0036] in: For a moment The remaining lifespan of the equipment, This is a weighting factor for the remaining lifespan of the equipment.

[0037] or

[0038] The dynamic reward function is:

[0039]

[0040] in: For a moment The predicted price of carbon trading market This represents the weighting coefficient for carbon trading market prices.

[0041] Preferably, in step S4, the optimization objective is... for:

[0042]

[0043] The range of the optimization objective is The smaller the value, the better the carbon emission optimization effect.

[0044] Preferably, in step S4, the deep reinforcement learning algorithm adopts a distributed training framework to accelerate the convergence of the model; and an experience replay mechanism is introduced during the training process to improve the generalization ability of the model.

[0045] Preferably, in step S4, the deep reinforcement learning algorithm is combined with a physical model, which includes hydraulic equations and equipment performance models; the combination method is to use the output of the physical model as a constraint condition for the deep reinforcement learning algorithm.

[0046] Preferably, in step S6, the dynamic adjustment of the pumped storage power station at the current moment... The operating parameters include pump speed, generator output power, and valve opening; a fault-tolerant mechanism is introduced during the adjustment process to ensure that basic operation can still be maintained in the event of equipment failure;

[0047] At the current moment When the operating parameters are dynamically adjusted, an adaptive control algorithm is introduced, which combines fuzzy control and sliding mode control to handle uncertainties and nonlinear problems.

[0048] The carbon optimization method for pumped storage power plants based on deep reinforcement learning provided by this invention has the following advantages:

[0049] By employing technologies such as multi-source data fusion, real-time monitoring, and intelligent prediction, this invention significantly improves the accuracy and efficiency of health monitoring for energy equipment. First, it integrates operational data from various sensors, including equipment status, environmental parameters, and historical fault records, enabling a comprehensive assessment of equipment health. Second, the system possesses real-time monitoring capabilities, allowing for rapid response to dynamic changes in equipment status and timely detection of potential faults, preventing equipment damage or downtime losses due to delayed monitoring. Furthermore, the intelligent prediction function based on big data analytics can identify complex fault modes in advance, providing a scientific basis for maintenance decisions, effectively extending equipment lifespan and reducing maintenance costs. This system also supports adaptive learning, automatically adjusting monitoring strategies according to changes in the equipment's operating environment to ensure efficient operation even under complex conditions. Overall, this invention provides a precise, real-time, and intelligent solution for health monitoring of energy equipment, significantly improving the operational reliability and economy of energy systems. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 The flowchart illustrates a carbon optimization method for pumped storage power plants based on deep reinforcement learning, as provided in this invention. Detailed Implementation

[0052] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0054] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0055] See Figure 1 This invention provides a carbon optimization method for pumped storage power plants based on deep reinforcement learning, comprising the following steps:

[0056] Step S1: Collect data on the pumped storage power station at the current time. The real-time operating data is collected and preprocessed to obtain preprocessed real-time operating data; wherein, the preprocessed real-time operating data includes the pumped storage power station at the current time. Water level, flow rate, equipment status and carbon emissions The real-time operational data also includes grid load demand, renewable energy generation, and carbon trading market prices.

[0057] In step S1, the real-time running data is preprocessed, including:

[0058] The real-time operating data is subjected to noise filtering to remove abnormal values ​​caused by sensor errors or environmental interference.

[0059] The real-time running data is preprocessed using a normalization function to obtain preprocessed real-time running data. The normalization method is as follows:

[0060] Read the running data of the most recent period T, and determine the minimum value among the running data of the most recent period T. and maximum value ; the current moment Real-time operational data is represented as The following formula is used to... Normalization is performed to obtain normalized real-time running data. ;

[0061]

[0062] The preprocessed real-time running data is stored in a temporary buffer.

[0063] Step S2: Invoke the multi-agent system; the multi-agent system is a cooperative system composed of multiple agents; each agent is assigned a specific optimization task; agents share state space and action space through a predefined communication protocol; each agent has a policy prediction model, and each agent generates the final policy through the policy prediction model.

[0064] In this step, the state space is defined as follows: Action space is defined as ;in, For the state space a state; In the action space A type of action.

[0065] The multi-agent system also includes a fault prediction module, which is used to predict potential equipment faults in advance; the fault prediction module makes predictions based on historical fault data and real-time operating data, using deep learning algorithms.

[0066] Step S3, Design the dynamic reward function The dynamic reward function With the current moment carbon emissions It is related to the operating efficiency of pumped storage power stations;

[0067] The dynamic reward function is:

[0068]

[0069] in:

[0070] , and The weighting coefficients are calculated using historical data normalization and optimized using a genetic algorithm.

[0071] For a moment The equipment efficiency of a pumped storage power station is given by the formula: ; )and , respectively time The device's output power and input power;

[0072] For a moment The economic indicators of pumped storage power stations are given by the following formula: ; and , respectively time The revenue and costs of pumped storage power stations;

[0073] For a moment The carbon emissions of pumped storage power stations are collected in real time through sensors;

[0074] or

[0075] The dynamic reward function is:

[0076]

[0077] in: For a moment The remaining lifespan of the equipment, This is a weighting factor for the remaining lifespan of the equipment.

[0078] or

[0079] The dynamic reward function is:

[0080]

[0081] in: For a moment The predicted price of carbon trading market This represents the weighting coefficient for carbon trading market prices.

[0082] Step S4: Using a deep reinforcement learning algorithm, with the optimization objective of achieving the best carbon emission performance based on the most recent period T, the policy prediction model of each agent in the multi-agent system is trained. During training, the agents select actions in the action space based on their current state and according to the dynamic reward function. Update the policy prediction model, and then make a decision on the current time based on the updated policy prediction model. Strategies;

[0083] In this step, the optimization objective is... for:

[0084]

[0085] The range of the optimization objective is The smaller the value, the better the carbon emission optimization effect.

[0086] In practical applications, the deep reinforcement learning algorithm adopts a distributed training framework to accelerate model convergence; an experience replay mechanism is introduced during the training process to improve the model's generalization ability.

[0087] In this step, the deep reinforcement learning algorithm is combined with a physical model, which includes hydraulic equations and equipment performance models; the combination method is to use the output of the physical model as the constraint condition of the deep reinforcement learning algorithm.

[0088] Step S5, for each agent at the current time A comprehensive analysis of the strategies is conducted to determine the current moment. A comprehensive strategy;

[0089] Step S6, based on the current time A comprehensive strategy to dynamically adjust the pumped storage power station at the current moment. The operating parameters are adjusted to minimize carbon emissions, and the current time after the operating parameters are adjusted is monitored in real time. carbon emissions The adjustment factor is obtained using the following formula. :

[0090]

[0091] in: For adjustment coefficients;

[0092] According to the adjustment factor The policy prediction model for each agent is dynamically adjusted to obtain the dynamically adjusted policy prediction model.

[0093] In this step, the dynamic adjustment of the pumped storage power station at the current moment... The operating parameters include pump speed, generator output power, and valve opening; a fault-tolerant mechanism is introduced during the adjustment process to ensure that basic operation can still be maintained in the event of equipment failure;

[0094] At the current moment When the operating parameters are dynamically adjusted, an adaptive control algorithm is introduced, which combines fuzzy control and sliding mode control to handle uncertainties and nonlinear problems.

[0095] Step S7, let Return to step S1.

[0096] The following is an example:

[0097] This embodiment provides a carbon optimization method for pumped storage power stations based on deep reinforcement learning. The method includes: collecting and preprocessing data to ensure data quality; constructing a multi-agent system to achieve distributed task processing and collaboration; designing a dynamic reward function that comprehensively considers multiple indicators; training the system using a deep reinforcement learning algorithm to improve its intelligence level; and applying optimization strategies and adjusting them in real time to ensure stable system operation. Through the synergistic effect of each step, this invention achieves low-carbon and efficient operation of pumped storage power stations, resulting in significant economic and environmental benefits.

[0098] The specific steps are as follows:

[0099] S1. Collect real-time operating data of pumped storage power stations, including water level, flow rate, equipment status and carbon emission data, and preprocess the data through a normalization function to remove noise and outliers. The preprocessed data is used as the input of the multi-agent system.

[0100] S2. Invoke the multi-agent system, where each agent is responsible for a different optimization task. The agents cooperate by sharing a state space and an action space. The state space is defined as follows: Action space is defined as ;

[0101] S3. Design a dynamic reward function to dynamically adjust the reward value based on real-time carbon emissions and power plant operating efficiency. The form of the reward function is as follows: ,

[0102] in, , and The weighting coefficients are calculated by normalizing historical data. For a moment The equipment efficiency is calculated using the following formula: , )and , respectively time The device's output power and input power; For a moment The economic indicators are calculated using the following formula: , For a moment Carbon emissions are collected in real time via sensors;

[0103] S4. Train a multi-agent system using a deep reinforcement learning algorithm (such as PPO) to optimize the policies of each agent. During training, the agents select actions based on their current state and update their policies according to the reward function. The optimization objective is... ,in, To optimize the time window, This represents the integration operation within the time window, and the range of the optimization objective is... The smaller the value, the better the carbon emission optimization effect;

[0104] S5. Apply the optimized strategy to the actual operation of the pumped storage power station, dynamically adjust operating parameters to minimize carbon emissions, and monitor carbon emission data in real time. Based on feedback, the strategy was further optimized, and the formula was adjusted as follows:

[0105]

[0106] in: For adjustment coefficients;

[0107] Furthermore, in step S1, real-time operational data, including water level, flow rate, equipment status, and carbon emission data, is first collected by sensors deployed at various key locations within the pumped storage power station. These sensors collect data at a high frequency and transmit it to the data processing module. Upon receiving the raw data, the data processing module first performs noise filtering to remove outliers caused by sensor errors or environmental interference. Then, a normalization function is applied... The data is normalized to ensure all data are within the same dimension, facilitating unified processing by the multi-agent system. The preprocessed data is stored in a temporary buffer, awaiting further analysis and processing. Step S1 achieves comprehensive collection and preprocessing of pumped-storage power station operation data. First, the collection of multi-dimensional data ensures the system's comprehensive understanding of the power station's operating status, providing a rich information foundation for subsequent optimization strategies. Second, noise filtering and normalization significantly improve data quality and usability, avoiding misjudgments due to data anomalies. Finally, the preprocessed data provides reliable input for the efficient operation of the multi-agent system, ensuring the stability and accuracy of the entire optimization process.

[0108] In step S2, a collaborative system consisting of multiple agents is constructed. Each agent is assigned a specific optimization task; for example, agent A is responsible for water pump scheduling, agent B focuses on improving power generation efficiency, and agent C monitors carbon emission control. The agents share a state space through a predefined communication protocol. and action space This architecture enables real-time information interaction and collaboration. Each agent integrates a prediction model based on historical data and an adjustment mechanism based on real-time data, allowing it to respond quickly and adjust its strategies in dynamic environments. Through step S2, a highly efficient collaborative multi-agent system is constructed. The multi-agent architecture achieves distributed task processing, significantly improving the system's parallel processing capabilities and response speed. The information sharing and collaboration mechanisms among agents ensure the achievement of the global optimization objective, avoiding the possibility of a single agent getting stuck in a local optimum. Furthermore, the prediction and adjustment mechanisms within the agents enhance the system's adaptability to dynamic environments, enabling it to maintain stable operation under complex conditions. Ultimately, this architecture provides strong technical support for the refined management of pumped storage power stations.

[0109] In step S3, a dynamically adjustable reward function was designed. This function comprehensively considers multiple key indicators such as equipment efficiency, economy, and carbon emissions, ensuring the comprehensiveness and balance of the optimization strategy. The dynamic adjustment mechanism enables the system to flexibly adjust its optimization direction based on real-time data, avoiding the strategy lag problem that may be caused by a static reward function. Ultimately, this design significantly improves the optimization performance of the multi-agent system, achieving the dual objectives of minimizing carbon emissions and maximizing economic efficiency.

[0110] Step S4 achieves efficient training and policy optimization for the multi-agent system. Deep reinforcement learning algorithms automatically learn the optimal policy, significantly improving the system's intelligence level. The experience replay mechanism enhances the model's generalization ability, enabling it to maintain stable performance under different operating conditions. The introduction of a distributed training framework greatly shortens training time and improves the system's practicality and scalability. Ultimately, this step ensures the scientific validity and effectiveness of the optimization strategy, providing a solid technical guarantee for the low-carbon operation of pumped storage power stations.

[0111] Step S5 enables the practical application and dynamic adjustment of the optimization strategy. Real-time parameter adjustment ensures the system can quickly respond to changes in operating status and always remain within the optimal operating range. The fault-tolerance mechanism enhances system reliability and avoids downtime losses due to equipment failure. Ultimately, this step transforms the theoretical optimization strategy into practical operational results, significantly improving the operating efficiency and carbon emission control level of the pumped storage power station.

[0112] Specifically, in S1, the real-time operating data also includes grid load demand, renewable energy generation, and carbon trading market prices;

[0113] The normalization function further includes normalization processing for grid load demand and renewable energy generation.

[0114] Furthermore, step S1 is extended to collect more types of real-time operational data, including grid load demand, renewable energy generation, and carbon trading market prices. This data is collected through additionally deployed sensors and data interfaces and transmitted to the data processing module along with existing data on water level, flow rate, equipment status, and carbon emissions. The normalization function not only processes the existing data but also normalizes the newly added grid load demand and renewable energy generation, ensuring all data are within the same dimension. The processed data is integrated into a comprehensive dataset, providing a more comprehensive input to the multi-agent system. By expanding the data collection scope, the system can more comprehensively perceive the operating environment and external conditions of the pumped storage power station. The addition of grid load demand and renewable energy generation enables the system to better predict and respond to dynamic changes in the grid, optimizing the power station's operating strategy to adapt to different load demands. The introduction of carbon trading market prices provides a reference for economic optimization, allowing the optimization strategy to not only focus on carbon emission reduction but also consider economic benefits. Ultimately, this extension significantly enhances the system's comprehensive optimization capabilities, making it perform better in complex and ever-changing operating environments.

[0115] Specifically, in S2, the multi-agent system realizes information sharing and cooperation among agents through a communication protocol;

[0116] The decision-making logic of the intelligent agent includes a prediction model based on historical data and an adjustment mechanism based on real-time data.

[0117] Furthermore, the multi-agent system in step S2 achieves information sharing and collaboration among agents through a communication protocol. Specifically, each agent not only handles its own tasks but also exchanges state information and decision results with other agents through a message passing mechanism. For example, agent A, responsible for water pump scheduling, sends the current state of the water pumps and the scheduling plan to agent B, responsible for improving power generation efficiency, so that B can adjust its power generation strategy based on the water pump's state. In addition, the agents' decision-making logic includes a prediction model based on historical data and an adjustment mechanism based on real-time data. The prediction model predicts future trends by analyzing historical data, while the adjustment mechanism corrects the prediction results based on real-time data. By introducing communication protocols and collaboration mechanisms, the multi-agent system achieves more efficient collaborative work. Information sharing allows each agent to understand the state and decisions of other agents, thereby making more coordinated decisions. The combination of prediction models and adjustment mechanisms improves the agents' adaptability to dynamic environments, enabling them to quickly adjust their strategies under changing conditions. Ultimately, this improvement significantly enhances the overall performance and response speed of the system, ensuring the real-time nature and accuracy of the optimization strategy.

[0118] Specifically, in S3, the weight coefficients of the dynamic reward function , and Optimization is performed using a genetic algorithm;

[0119] The dynamic reward function also includes consideration of equipment lifespan, and takes the following form:

[0120]

[0121] in, For the remaining lifespan of the equipment, These are the weighting coefficients.

[0122] Furthermore, by optimizing the weight coefficients using a genetic algorithm, the dynamic reward function can more accurately reflect the system's multi-objective optimization needs. The inclusion of equipment lifespan ensures that while optimizing carbon emissions and economic efficiency, the system also considers the long-term health of the equipment, preventing premature damage due to over-optimization. Ultimately, this improvement not only enhances the scientific rigor of the optimization strategy but also extends equipment lifespan and reduces maintenance costs.

[0123] Specifically, in S4, the deep reinforcement learning algorithm employs a distributed training framework to accelerate model convergence;

[0124] An experience replay mechanism is introduced during the training process to improve the model's generalization ability.

[0125] Furthermore, the deep reinforcement learning algorithm in step S4 employs a distributed training framework to accelerate model convergence. Specifically, the training process is decomposed into multiple subtasks, distributed to different computing nodes for parallel processing. Each node independently trains its own model and periodically exchanges model parameters with other nodes, integrating them through a parameter server. In addition, an experience replay mechanism is introduced during training, learning from randomly sampled historical data to avoid the influence of data order on model training, thus improving the model's generalization ability. The distributed training framework significantly improves the model's training efficiency, enabling the system to process large-scale datasets in a shorter time. The experience replay mechanism enhances the model's generalization ability, allowing it to maintain stable performance under different operating conditions. Ultimately, this improvement not only accelerates model convergence but also enhances the system's practicality and scalability, enabling it to better adapt to complex real-world applications.

[0126] Specifically, in S5, the dynamically adjusted operating parameters include pump speed, generator output power, and valve opening.

[0127] The adjustment process incorporates a fault-tolerant mechanism to ensure that basic operation can be maintained even in the event of equipment failure.

[0128] Furthermore, step S5 involves dynamically adjusting operating parameters including pump speed, generator output power, and valve opening. Specifically, the system dynamically adjusts these parameters based on real-time data and optimization strategies using a control algorithm. For example, when an increase in carbon emissions is detected, the system reduces pump speed to decrease energy consumption and adjusts generator output power to maintain grid stability. In addition, a fault-tolerant mechanism is introduced during the adjustment process. When a device fails, the system can automatically switch to backup equipment or adjust the operating mode to ensure the basic operation of the power station remains unaffected. By dynamically adjusting operating parameters, the system can quickly respond to changes in operating status and always remain within the optimal operating range. This fault-tolerant mechanism enhances system reliability and avoids downtime losses due to equipment failure. Ultimately, this improvement significantly enhances the system's operational stability and reliability, ensuring efficient operation of the pumped storage power station under various operating conditions.

[0129] Specifically, in S3, the dynamic reward function also includes a prediction of carbon trading market prices, in the form of:

[0130]

[0131] in, This is a price forecast for the carbon trading market. These are the weighting coefficients.

[0132] Furthermore, by incorporating carbon trading market price forecasts, the system can more comprehensively consider economic factors. The optimization strategy not only focuses on reducing carbon emissions but also adjusts to market price fluctuations to achieve greater economic benefits. Ultimately, this improvement enables the system to achieve environmental goals while also generating greater economic returns for the power plant.

[0133] Specifically, in S2, the multi-agent system also includes a fault prediction module for detecting potential equipment faults in advance;

[0134] The fault prediction module uses a deep learning algorithm to make predictions based on historical fault data and real-time operational data.

[0135] Furthermore, a fault prediction module is added to the multi-agent system in step S2. Specifically, the fault prediction module analyzes historical fault data and real-time operating data, using deep learning algorithms (such as LSTM) to predict potential equipment faults. For example, the module analyzes the equipment's vibration data, temperature data, and operating time to identify possible fault modes and issue early warnings. These warnings are then transmitted to the corresponding agents so they can adjust their strategies to prevent faults from impacting system operation. The addition of the fault prediction module significantly improves the system's preventative maintenance capabilities, enabling early detection of potential faults and the implementation of corrective measures to avoid downtime and economic losses caused by sudden failures. Ultimately, this improvement not only enhances system reliability but also extends equipment lifespan and reduces maintenance costs.

[0136] Specifically, in S4, the deep reinforcement learning algorithm is combined with a physical model, which includes hydraulic equations and equipment performance models.

[0137] The combination method involves using the output of the physical model as a constraint condition for the deep reinforcement learning model.

[0138] Furthermore, step S4 integrates the deep reinforcement learning algorithm with a physical model. Specifically, the physical model includes hydraulic equations and equipment performance models, whose outputs serve as constraints for the deep reinforcement learning model. For example, the hydraulic equations are used to calculate the dynamic changes in water flow, and the equipment performance model is used to predict the efficiency of equipment under different operating conditions. During training, the deep reinforcement learning model not only maximizes the reward function but also satisfies these physical constraints, ensuring the feasibility and safety of the optimization strategy. By integrating the physical model, the deep reinforcement learning algorithm can more accurately simulate the actual operating environment, and the optimization strategy is more in line with physical laws and equipment performance limitations. Ultimately, this improvement enhances the reliability and safety of the optimization strategy, ensuring the stable operation of the system in practical applications.

[0139] Specifically, in S5, the real-time adjustment process introduces an adaptive control algorithm to dynamically adjust the control parameters based on real-time data;

[0140] The adaptive control algorithm combines fuzzy control and sliding mode control to handle uncertainties and nonlinear problems.

[0141] Furthermore, the real-time adjustment process in step S5 incorporates an adaptive control algorithm. Specifically, the system employs a combination of fuzzy control and sliding mode control to handle uncertainties and nonlinearities during operation. For example, when a small change in equipment parameters is detected, fuzzy control can quickly adjust the control parameters, while sliding mode control ensures the system remains stable during changes. This combined algorithm allows the system to better adapt to complex dynamic environments. The introduction of the adaptive control algorithm significantly improves the system's adaptability and stability, enabling it to operate efficiently in the face of uncertainties and nonlinearities. Ultimately, this improvement not only enhances the system's robustness but also strengthens its optimization performance under complex operating conditions.

[0142] This embodiment also provides a computer device applicable to the carbon optimization method for pumped storage power stations based on deep reinforcement learning, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the carbon optimization method for pumped storage power stations based on deep reinforcement learning as proposed in the above embodiment.

[0143] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0144] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the carbon optimization method for pumped storage power stations based on deep reinforcement learning as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0145] In summary, this invention significantly improves the accuracy and efficiency of energy equipment health monitoring through multi-source data fusion, real-time monitoring, and intelligent prediction. First, the method integrates operational data from different sensors, including equipment status, environmental parameters, and historical fault records, enabling a comprehensive assessment of equipment health. Second, the system possesses real-time monitoring capabilities, rapidly responding to dynamic changes in equipment status and promptly identifying potential faults, thus preventing equipment damage or downtime losses due to delayed monitoring. Furthermore, the intelligent prediction function based on big data analysis can identify complex fault modes in advance, providing a scientific basis for maintenance decisions, effectively extending equipment lifespan and reducing maintenance costs. This system also supports adaptive learning, automatically adjusting monitoring strategies according to changes in the equipment's operating environment, ensuring efficient operation even under complex conditions. Overall, this invention provides a precise, real-time, and intelligent solution for energy equipment health monitoring, significantly improving the operational reliability and economy of energy systems.

[0146] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A carbon optimization method for pumped storage power plants based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Collect data on the pumped storage power station at the current time. The real-time operating data is collected and preprocessed to obtain preprocessed real-time operating data; wherein, the preprocessed real-time operating data includes the pumped storage power station at the current time. Water level, flow rate, equipment status and carbon emissions ; Step S2: Invoke the multi-agent system; the multi-agent system is a cooperative system composed of multiple agents; each agent is assigned a specific optimization task; agents share state space and action space through a predefined communication protocol; each agent has a policy prediction model, and each agent generates the final policy through the policy prediction model. Step S3, Design the dynamic reward function The dynamic reward function With the current moment carbon emissions It is related to the operating efficiency of pumped storage power stations; The dynamic reward function is: ; in: , and The weighting coefficients are calculated using historical data normalization and optimized using a genetic algorithm. For a moment The equipment efficiency of a pumped storage power station is given by the formula: ; )and , respectively time The device's output power and input power; For a moment The economic indicators of pumped storage power stations are given by the following formula: ; and , respectively time The revenue and costs of pumped storage power stations; For a moment The carbon emissions of pumped storage power stations are collected in real time through sensors; or The dynamic reward function is: ; in: For a moment The remaining lifespan of the equipment, This is a weighting factor for the remaining lifespan of the equipment. or The dynamic reward function is: ; in: For a moment The predicted price of carbon trading market This represents the weighting coefficient for carbon trading market prices; Step S4: Using a deep reinforcement learning algorithm, with the optimization objective of achieving the best carbon emission performance based on the most recent period T, the policy prediction model of each agent in the multi-agent system is trained. During training, the agents select actions in the action space based on their current state and according to the dynamic reward function. Update the policy prediction model, and then make a decision on the current time based on the updated policy prediction model. Strategies; Step S5, for each agent at the current time A comprehensive analysis of the strategies is conducted to determine the current moment. A comprehensive strategy; Step S6, based on the current time A comprehensive strategy to dynamically adjust the pumped storage power station at the current moment. The operating parameters are adjusted to minimize carbon emissions, and the current time after the operating parameters are adjusted is monitored in real time. carbon emissions The adjustment factor is obtained using the following formula. : ; in: For adjustment coefficients; According to the adjustment factor The policy prediction model for each agent is dynamically adjusted to obtain the dynamically adjusted policy prediction model. Step S7, let Return to step S1.

2. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, The real-time operational data also includes grid load demand, renewable energy generation, and carbon trading market prices.

3. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S1, the real-time running data is preprocessed, including: The real-time operating data is subjected to noise filtering to remove abnormal values ​​caused by sensor errors or environmental interference; The real-time running data is preprocessed using a normalization function to obtain preprocessed real-time running data. The normalization method is as follows: Read the running data for the most recent period T, and determine the minimum value among the running data for the most recent period T. and maximum value ; the current moment Real-time operational data is represented as The following formula is used to... Normalization is performed to obtain normalized real-time running data. ; ; The preprocessed real-time running data is stored in a temporary buffer.

4. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S2, the state space is defined as follows: Action space is defined as ;in, For the state space a state; In the action space A type of action.

5. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, The multi-agent system also includes a fault prediction module, which is used to predict potential equipment faults in advance; the fault prediction module makes predictions based on historical fault data and real-time operating data, using deep learning algorithms.

6. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the optimization objective is... for: ; The range of the optimization objective is The smaller the value, the better the carbon emission optimization effect.

7. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the deep reinforcement learning algorithm adopts a distributed training framework to accelerate the convergence of the model; an experience replay mechanism is introduced during the training process to improve the generalization ability of the model.

8. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the deep reinforcement learning algorithm is combined with a physical model, which includes hydraulic equations and equipment performance models; the combination method is to use the output of the physical model as a constraint condition for the deep reinforcement learning algorithm.

9. The carbon optimization method for pumped storage power stations based on deep reinforcement learning according to claim 1, characterized in that, In step S6, the dynamic adjustment of the pumped storage power station at the current moment... The operating parameters include pump speed, generator output power, and valve opening; a fault-tolerant mechanism is introduced during the adjustment process to ensure that basic operation can still be maintained in the event of equipment failure; At the current moment When the operating parameters are dynamically adjusted, an adaptive control algorithm is introduced, which combines fuzzy control and sliding mode control to handle uncertainties and nonlinear problems.

Citation Information

Patent Citations

  • Environment-friendly micro-grid optimization scheduling method and system based on deep reinforcement learning

    CN117726143A

  • Power distribution network source network load storage low-carbon optimization scheduling increment reinforcement learning method and system

    CN119623567A