Integrated energy system low-carbon scheduling method and system based on carbon capture and demand response
By applying reinforcement learning algorithms to optimize low-carbon scheduling, the problems of high carbon emissions of thermal power generation and intermittent clean energy are solved, and the system's low-carbon and efficient operation and dynamic adaptability are achieved.
Patent Information
- Application Number
- CN202510062928.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-27
AI Technical Summary
At present, thermal power generation accounts for a large proportion in the energy structure, and the high carbon emission characteristics are inconsistent with the low carbon development demand. Traditional carbon capture efficiency is low and energy consumption is high, making it difficult to meet the demand for large-scale emission reduction. At the same time, the intermittent and volatility problems of clean energy such as wind power and photovoltaics are difficult to effectively solve.
The reinforcement learning algorithm based on experience replay strategy-value network is adopted to optimize the integrated energy system low-carbon scheduling. Through multi-agent division, each subsystem is divided into independent optimization units, scheduling actions are generated and reward values are calculated based on the execution situation, and the policy network parameters are continuously optimized to achieve low-carbon scheduling.
It realizes low-carbon and efficient operation of the integrated energy system, dynamically adapts to environmental changes, improves wind and light generation utilization, reduces wind and light abandonment phenomena, optimizes the power generation and carbon capture efficiency of thermal power units, maximizes the utilization of P2G equipment, and reduces peak load through user demand response.
Smart Images

Figure CN120046904A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of thermal power generation, and particularly to a low-carbon scheduling method, device, system, computer device, and computer-readable storage medium for an integrated energy system based on carbon capture and demand response. Background Art
[0002] In response to China's "dual carbon" goal, the integrated energy system, under the integration of various energies such as cold, heat, electricity, and gas, has gradually become a key technology for improving energy utilization efficiency and reducing carbon emissions.
[0003] Since the current thermal power still accounts for a large proportion in the energy structure, there is a contradiction between its high carbon emission characteristics and the low-carbon development requirements. Although carbon capture technology can reduce the carbon emissions of thermal power, traditional carbon capture has low efficiency and high energy consumption under high load, and it is difficult to meet the large-scale emission reduction requirements. At the same time, the installed capacity of clean energies such as wind power and photovoltaic power has increased, making the problems of intermittency and volatility of power supply increasingly prominent, and the existing peak shaving means such as energy storage are difficult to support the dynamic requirements of the system in terms of capacity and cost. Summary of the Invention
[0004] Embodiments of the present application provide a low-carbon scheduling method, system, computer device, and computer-readable storage medium for an integrated energy system based on carbon capture and demand response, so as to at least solve the problem of energy utilization efficiency of the integrated energy system in related technologies.
[0005] In a first aspect, embodiments of the present application provide a low-carbon scheduling method for an integrated energy system based on carbon capture and demand response, characterized in that the low-carbon scheduling optimization of the integrated energy system is performed through a reinforcement learning algorithm of a policy-value network based on experience replay. The method includes:
[0006] Establish a system model of the integrated energy system, where the integrated energy system includes: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem, and a user demand response subsystem, and a carbon capture device is configured in the thermal power subsystem;
[0007] Perform preprocessing on the system model, including: for each subsystem of the integrated energy system, define the state space, action space, and reward function corresponding to the reinforcement learning algorithm, and divide each subsystem into independent optimization units through multi-agent partitioning;
[0008] Through the reinforcement learning algorithm, perform low-carbon scheduling optimization based on the system model after the preprocessing.
[0009] In some of these embodiments, performing low-carbon scheduling optimization based on the system model after the preprocessing includes:
[0010] The agents of each subsystem read the status information related to the tasks of each subsystem;
[0011] According to the status information, generate a scheduling action through the policy network, and feedback the scheduling action to the system model;
[0012] Respectively instruct each subsystem to respond to the scheduling action, and after the execution of the scheduling action is completed, through the value network, calculate the reward value according to the execution situation, with the total system operating cost, carbon emission and wind-solar consumption rate as the measurement indicators;
[0013] Obtain historical data from the experience replay pool, continuously optimize the parameters of the policy network according to the historical data and the reward value, and generate a scheduling plan for low-carbon scheduling through the continuously optimized system model.
[0014] In some embodiments, the state space includes the system global state and the local states of each subsystem, where:
[0015] The system global state includes: wind-solar power generation prediction value, real-time grid load demand, total carbon emissions and carbon trading price;
[0016] The local states of each subsystem include: wind power generation data, photovoltaic power generation data, wind curtailment data and light curtailment data, thermal power unit power output data, carbon capture equipment operation data, operating power and gas production of the energy conversion subsystem, and user demand response situation.
[0017] In some embodiments, defining the action space for each subsystem includes:
[0018] For the agents of the wind power subsystem and the photovoltaic subsystem: formulate a priority power distribution plan for power generation. When the grid demand load is greater than or equal to the wind-solar power generation, all the wind-solar power is used for the grid load;
[0019] When the grid demand load is less than the wind-solar power generation, of the wind-solar power generation is used to meet the grid load demand, and the power beyond meeting the grid load demand is used to drive the energy conversion subsystem;
[0020] For the agents of the thermal power subsystem: by adjusting its output power and carbon capture rate, when the wind-solar power generation cannot meet the grid load demand, supplement the gap of the grid demand load through the thermal power subsystem;
[0021] When the power generation of the thermal power unit can meet the gap of the demand load of the power grid and there is surplus power, the surplus power is used to drive the carbon capture device to capture carbon dioxide in the flue gas discharged by the thermal power subsystem;
[0022] For the agent of the energy conversion subsystem: when the wind and solar power generation is excessive, through the energy conversion subsystem, the excessive power is used to produce hydrogen, and the hydrogen reacts with the carbon dioxide in the carbon capture device to generate methane gas;
[0023] For the agent of the user demand response subsystem: through preset incentive measures, users are guided to adjust their electricity consumption behaviors to reduce the electricity load during peak hours or transfer the electricity demand to off-peak hours.
[0024] In some embodiments, the action space of the agent of the user demand response subsystem includes: an incentive intensity dimension, a user type dimension, and a response time dimension, where:
[0025] The incentive intensity dimension is used to determine the price discount or reward intensity provided to users, including low incentive level, medium incentive level, and high incentive level;
[0026] The user type dimension is used to classify users into residential users, industrial users, and commercial users based on the characteristics of different users, and corresponding action strategies are set according to different user types;
[0027] The response time dimension is used to select the response time for the incentive effect according to the operating state of the power grid, and the response time includes peak hours, off-peak hours, and load balancing hours.
[0028] In some embodiments, the defined reward function includes:
[0029] Weighted summation is performed according to carbon emissions, operating costs, wind and solar power consumption rates, and load balancing parameters to obtain the global reward;
[0030] The reward values and penalty values are respectively determined according to the grid connection ratios and curtailment ratios of the wind power subsystem and the photovoltaic subsystem, and the wind power reward and the photovoltaic power reward are respectively obtained by weighted summation of the reward values and the penalty values;
[0031] Multiple thermal power reward and penalty values are respectively determined according to the load supplement amount, carbon capture amount, and fuel usage amount of the thermal power unit, and the thermal power unit reward is obtained by weighted summation of each thermal power reward and penalty value;
[0032] Determine multiple conversion reward values according to the curtailed wind and photovoltaic electric energy absorbed by the energy conversion subsystem and the methane conversion efficiency, and obtain the energy conversion equipment reward by weighted summation of each conversion reward value;
[0033] Determine the user demand reward value and the user demand penalty value respectively according to the actual reduced electricity load and incentive expenditure of the user, and obtain the user demand response reward by weighted summation of the user demand reward value and the user demand penalty value.
[0034] In some embodiments, the update process of the policy network parameters includes:
[0035] Packet store the historical feedback data obtained by the integrated energy system executing the feedback action to generate an experience replay pool, where the historical feedback data includes the current state of the agent, the executed action, the obtained reward, the new state after executing the action, and the parameters of the historical policy network;
[0036] Randomly extract target data from the experience replay pool, and based on the target data, obtain the probability value when the new policy network and the old policy network select the same action, and use the ratio of the probability values as the calculation weight for updating the policy network in the next step;
[0037] Subtract the overall reward value of the next state from the overall reward value of the current state to obtain the calculation step size for updating the policy network in the next step;
[0038] Based on the calculation weight and the calculation step size, use the gradient descent method to update the parameters of the policy network to obtain a new policy network.
[0039] In a second aspect, an embodiment of the present application provides a low-carbon scheduling system for an integrated energy system based on carbon capture and demand response, which is characterized in that the integrated energy system is optimized for low-carbon scheduling through a reinforcement learning algorithm of a policy-value network based on experience replay. The system includes: a model establishment module, a preprocessing module, and a scheduling optimization module, where:
[0040] The model establishment module is used to establish a system model of the integrated energy system, where the integrated energy system includes: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem, and a user demand response subsystem, and a carbon capture device is configured in the thermal power subsystem;
[0041] The preprocessing module is used to preprocess the system model, including: defining the state space, action space, and reward function corresponding to the reinforcement learning algorithm for each subsystem of the integrated energy system, and dividing each subsystem into independent optimization units through multi-agent partitioning;
[0042] The scheduling optimization module is configured to perform low-carbon scheduling optimization based on the preprocessed system model through the reinforcement learning algorithm.
[0043] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect above is implemented.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect above is implemented.
[0045] Compared with the related art, the low-carbon scheduling method for an integrated energy system based on carbon capture and demand response provided by the embodiment of the present application adjusts scheduling actions in a timely manner according to continuously changing state information to adapt to various working conditions. In each decision-making cycle, through the loop of generating actions, executing actions, calculating rewards, and optimizing parameters by the policy network, the system always operates in the direction of low-carbon, economic, and efficient, can dynamically adapt to environmental changes (such as fluctuations in wind and solar power generation, changes in grid load demand, user behavior responses, etc.), continuously optimize the scheduling strategies of each subsystem, effectively improve the utilization rate of wind and solar power generation, reduce the phenomenon of wind and light abandonment, optimize the power generation and carbon capture efficiency of thermal power units, maximize the utilization of excess wind and solar power by P2G devices, and at the same time reduce peak loads through user demand response, and overall achieve the low-carbon and efficient operation of the integrated energy system. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0047] Figure 1 is a flowchart of a low-carbon scheduling method for an integrated energy system based on carbon capture and demand response according to an embodiment of the present application;
[0048] Figure 2 is a schematic structural diagram of an integrated energy system according to an embodiment of the present application;
[0049] Figure 3 is a schematic diagram of a carbon capture device according to an embodiment of the present application;
[0050] Figure 4 is a structural block diagram of a low-carbon scheduling system for an integrated energy system based on carbon capture and demand response according to an embodiment of the present application;
[0051] Figure 5It is a diagram of the combined operation mode of a low-carbon integrated energy system according to an embodiment of the present application;
[0052] Figure 6 It is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0053] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application.
[0054] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar scenarios based on these drawings without creative efforts. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0055] When "embodiment" is mentioned in the present application, it means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0056] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "coupled" and "linked" involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0057] In response to China's "dual carbon" goal, under the integration of multiple energy sources such as cold, heat, electricity, and gas, the integrated energy system has gradually become a key technology for improving energy utilization efficiency and reducing carbon emissions. However, due to the still large proportion of thermal power in the energy structure, there is a contradiction between its high carbon emission characteristics and the low-carbon development requirements. Although carbon capture technology can reduce carbon emissions from thermal power, traditional carbon capture has low efficiency and high energy consumption under high loads and is difficult to meet the demand for large-scale emission reduction.
[0058] At the same time, the installed capacity of clean energy such as wind power and photovoltaic power has increased, making the problems of intermittency and volatility of power supply increasingly prominent. Existing peak shaving means such as energy storage are difficult to support the dynamic demand of the system in terms of capacity and cost. Although demand response technology can shave peaks and fill valleys on the load side, it is mostly limited to a single power system and is difficult to achieve collaborative optimization in a multi-energy scenario, resulting in a low consumption rate of clean energy. In addition, the carbon trading mechanism still mainly focuses on fixed prices and fails to encourage enterprises to actively reduce emissions.
[0059] In view of this, an embodiment of the present application provides a low-carbon scheduling method for an integrated energy system combining carbon capture and demand response. This method first constructs a multi-energy system architecture covering cooling, heating, electricity, and gas, including main energy supply devices such as thermal power plants, wind power plants, and photovoltaic power plants, as well as energy conversion devices such as power-to-gas (P2G) devices, to achieve the conversion, storage, and output of different energy forms. On this basis, mathematical models of each subsystem are established in detail, such as the model of a thermal power unit with carbon capture, the model of a wind turbine generator, the model of a photovoltaic generator, and the chemical reaction and energy conversion model of a P2G device, to support the dynamic response of the system.
[0060] Figure 1 is a flowchart of a low-carbon scheduling method for an integrated energy system based on carbon capture and demand response according to an embodiment of the present application, as Figure 1 shown, and this process includes the following steps:
[0061] S101, establish a system model of the integrated energy system, where the integrated energy system includes: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem, and a user demand response subsystem, and among them, the thermal power subsystem includes a carbon capture device;
[0062] Figure 2 is a schematic structural diagram of the integrated energy system according to an embodiment of the present application, as Figure 2 shown, the thermal power subsystem, as an important energy supply part of the system, plays a key role in energy supply. It includes a carbon capture device, and this design aims to address the problem of high carbon emissions from thermal power.
[0063] In this embodiment, the thermal power subsystem is configured with a carbon capture device. Through the carbon capture device, carbon dioxide in the flue gas can be captured during the power generation process, reducing carbon emissions and enabling it to develop towards a low-carbon direction while meeting energy demands. For example, during the operation of thermal power, the carbon capture device can effectively capture a certain proportion of carbon dioxide according to system requirements and operating parameters, reducing the impact on the environment.
[0064] Furthermore, the wind power subsystem and the photovoltaic subsystem belong to the clean energy supply subsystems, which are important components for realizing the green transformation of energy. They can convert renewable wind energy and solar energy into electricity and provide clean power for the system.
[0065] In addition, the energy conversion subsystem is mainly used to achieve the conversion of different energy forms. For example, a power-to-gas (P2G) device can convert excess electric energy (such as when wind and solar power generation is excessive) into natural gas. It can not only effectively utilize excess energy but also achieve energy storage and cross-form conversion through chemical reactions (such as electrolyzing water to produce hydrogen and then reacting with carbon dioxide to produce methane), improving the comprehensive energy utilization efficiency and enhancing the flexibility and stability of the system.
[0066] Finally, the user demand response subsystem guides users to adjust their electricity consumption behaviors by formulating incentive measures. According to the characteristics of different user groups (residential, industrial, commercial) and the operating status of the power grid (peak, valley, load balancing periods), it provides schemes with different incentive intensities (low, medium, high) to encourage users to reduce load during peak periods and increase load during valley periods, achieving peak shaving and valley filling, improving the system's load balancing ability, reducing energy waste, and also helping to enhance users' participation and interactivity in the energy system.
[0067] S102, preprocess the system model, including: for each subsystem of the integrated energy system, define the state space, action space, and reward function corresponding to the reinforcement learning algorithm, and divide each subsystem into independent optimization units through multi-agent partitioning;
[0068] Among them, define the state space, action space, and reward function of the reinforcement learning algorithm for each subsystem of the integrated energy system. These parameters are the basis for constructing an intelligent control system, enabling each subsystem to make decisions based on its own state and continuously optimize its behavior through reward feedback.
[0069] In this embodiment, the system is divided into multiple independent units, and each unit of each subsystem is controlled by an agent. This distributed architecture is beneficial to improving the flexibility and adaptability of the system. Each agent can process tasks in parallel, make independent decisions according to the characteristics and state of the subsystem it is responsible for, and at the same time can achieve the optimization of the entire system through interaction and cooperation with other agents.
[0070] Specifically, the state space includes the global state of the system and the local states of each subsystem, where:
[0071] The global state of the system includes: predicted values of wind and solar power generation, real-time load demand of the power grid, total carbon emissions, and carbon trading price;
[0072] The local states of each subsystem include: wind power generation data, photovoltaic power generation data, wind curtailment data and PV curtailment data, power output data of thermal power units, operation data of carbon capture equipment, operating power and gas production of the energy conversion subsystem, and user demand response situation.
[0073] Furthermore, define the action space for each subsystem to include:
[0074] (1) For the agents of the wind power subsystem and the photovoltaic subsystem: formulate a priority power distribution plan for power generation. When the grid demand load is greater than or equal to the wind and solar power generation, all the wind and solar power is used for the grid load;
[0075] When the demand load of the power grid is less than the wind and solar power generation, part of the wind and solar power generation is used to meet the power grid load demand, and the electricity beyond that for meeting the power grid load demand is used to drive the energy conversion subsystem.
[0076] It can be understood that in this embodiment, for the intelligent agent of wind and solar power generation, the design of its action space aims to make full use of wind and solar power generation resources and achieve efficient distribution of electric energy. When the power grid load demand is large, power supply to the power grid is prioritized to ensure stable power supply. When there is an oversupply of wind and solar power generation, the excess electricity is used to drive power-to-gas (P2G) equipment to realize energy storage and conversion, improve the comprehensive energy utilization efficiency, reduce the phenomenon of curtailment of wind and solar power, and enhance the consumption capacity of clean energy in the system.
[0077] (2) For the intelligent agent of the thermal power subsystem: By adjusting its output power and carbon capture rate, when the power generated by wind and solar cannot meet the power grid load demand, the thermal power unit is used to supplement the gap in the power grid demand load.
[0078] When the power generation of the thermal power unit can meet the gap in the power grid demand load and there is surplus electricity, the surplus electricity is used to drive the carbon capture device to capture carbon dioxide in the flue gas emitted by the thermal power unit.
[0079] Among them, Figure 3 is a schematic diagram of a carbon capture device according to an embodiment of the present application. As Figure 3 shown, when the power generated by the thermal power unit can meet the load gap and there is surplus electricity, the carbon capture device installed in the flue gas pipeline of the thermal power unit will be activated, and the surplus electricity is used to drive the carbon capture device to capture carbon dioxide in the flue gas.
[0080] For the thermal power subsystem, in this embodiment, its action space focuses on supplementing the power grid load gap and carbon capture. When the wind and solar power generation is insufficient, the output power is adjusted in a timely manner to fill the power gap and ensure the power supply reliability of the system. When there is surplus electricity, the carbon capture device is activated to capture carbon dioxide using the excess electric energy, which not only reduces carbon emissions but also realizes the coordinated optimization of thermal power operation and carbon emission reduction, promoting the system to develop towards low carbon while meeting energy demands.
[0081] (3) For the intelligent agent of the energy conversion subsystem: When there is an oversupply of wind and solar power generation, through the energy conversion subsystem, the excess electricity is used to produce hydrogen, and the hydrogen reacts with carbon dioxide in the carbon capture device to generate methane gas.
[0082] Optionally, the energy conversion subsystem can be a power-to-gas (P2G) device. It can be understood that by starting the P2G device, unstable electric energy is converted into storable methane gas, while consuming the carbon dioxide captured by the thermal power unit, forming a mechanism for energy conversion and recycling. This not only improves the system's ability to accommodate clean energy but also enhances energy storage and regulation capabilities, contributing to coping with the dynamic changes in energy supply and demand and improving the overall stability and flexibility of the system.
[0083] (3) For the agents in the user demand response subsystem: By presetting incentive measures, users are guided to adjust their electricity consumption behaviors to reduce the electricity load during peak hours or shift the electricity demand to off-peak hours.
[0084] In this embodiment, by formulating incentive measures and based on the characteristics of different user groups and the operating status of the power grid, users are guided to adjust their electricity consumption behaviors. During peak hours, the electricity load is reduced to relieve the pressure on the power grid and avoid power supply shortages; during off-peak hours, the electricity demand is increased to improve the operating efficiency of the power grid and achieve peak shaving and valley filling. This helps optimize the system load curve, improve energy utilization efficiency, enhance users' participation and interactivity in the energy system, and promote the balance between supply and demand.
[0085] More specifically, the action space of the household demand response agent is divided into three key dimensions: incentive intensity, user grouping, and response period.
[0086] Among them, the incentive intensity is divided into three levels: low, medium, and high. This grading method can implement differential incentives for user groups with different price sensitivities. For users who are not sensitive to price changes, the low-incentive plan can guide their participation in demand response without excessive cost increase; while medium and high incentives are targeted at price-sensitive users, attracting them to actively reduce the load through higher rewards, fully tapping the adjustment potential on the user side, and improving the effect of demand response.
[0087] In addition, users are divided into residential, industrial, and commercial users according to their electricity consumption behaviors, taking into account the characteristics of different user types. Residential users have limited ability to reduce the load under comfort requirements, so incentive measures need to consider their living needs; industrial users have great potential to reduce the load but are affected by production, and the pros and cons need to be weighed; commercial users are sensitive to incentives and flexible, and can quickly adjust their electricity consumption behaviors through reasonable incentives. Such grouping helps formulate incentive strategies that better suit the actual situation of users, improve user participation, and the effectiveness of responses.
[0088] In this embodiment, peak, valley, and load balancing periods are selected according to the grid operation status for incentive effects, realizing the dynamic matching of user needs and grid supply and demand. Reducing electricity demand during peak periods can relieve the grid pressure and avoid overload; encouraging increased electricity consumption during valley periods can improve grid efficiency and promote the rational use of energy; and load balancing periods can maintain the stable operation of the grid. This precise incentive according to periods enhances the grid's load regulation ability and improves the overall operation efficiency of the system.
[0089] It can be understood that the solution of this application is divided into three key dimensions: incentive intensity, user clustering, and response period through the demand response mechanism, and a response strategy for peak shaving and valley filling is designed for the electrical load to reduce carbon emissions; these three dimensions cooperate with each other to encourage users to adjust their electricity consumption behaviors from different perspectives, improving the enthusiasm and initiative of users to participate in demand response. Different user groups can flexibly choose participation methods according to their own situations under different incentive intensities and response periods, increasing the flexibility and diversity of demand response, enabling the system to better adapt to various operating conditions.
[0090] In addition, defining the reward function includes the following steps:
[0091] Step1, perform weighted summation according to carbon emissions, operating costs, wind and solar energy consumption rates, and load balancing parameters to obtain the global reward;
[0092] In this embodiment, carbon emissions, operating costs, wind and solar energy consumption rates, and load balancing are used as core indicators to comprehensively measure the comprehensive performance of the integrated energy system. Through the comprehensive consideration of these key indicators, the system is promoted to pursue low carbon emissions, low-cost operation, efficient use of clean energy, and stable supply and demand balance during operation.
[0093] In addition, it should be noted that the specific method of weighted summation can be flexibly adjusted according to the system's emphasis on each indicator, guiding the system to develop towards the overall optimal direction and achieving multi-objective collaborative optimization.
[0094] Step2, according to the grid connection ratios and curtailment ratios of the wind power subsystem and the photovoltaic subsystem, determine the reward values and penalty values respectively, and perform weighted summation of the reward values and penalty values to obtain the wind power generation reward and the photovoltaic power generation reward respectively;
[0095] It can be understood that the wind and solar power generation agents are encouraged to give priority to meeting the grid power demand, increasing the proportion of clean energy in the energy supply. At the same time, the phenomenon of wind and light curtailment is punished, prompting the system to actively seek ways to consume excess wind and solar power, such as through power-to-gas (P2G) equipment conversion, etc., thereby improving the utilization rate of wind and solar power generation resources, reducing energy waste, and enhancing the stability and reliability of clean energy in the system.
[0096] Step 3: Determine multiple reward and penalty values respectively based on the load supplement amount, carbon capture amount, and fuel usage amount of the thermal power unit, and sum up the weighted values of each reward and penalty value to obtain the reward for the thermal power unit.
[0097] Specifically, for the dispatching optimization method of this application, on the one hand, it rewards the thermal power unit for effectively supplementing the power grid load gap to ensure the stability and reliability of system power supply. On the other hand, it encourages increasing the carbon dioxide capture amount to promote the low-carbon development of thermal power units and reduce the environmental impact of carbon emissions. At the same time, considering the fuel usage amount and the operating cost of carbon capture equipment, it prompts the thermal power unit to optimize the energy utilization efficiency during operation, reduce the operating cost, and achieve the balanced development of power generation and emission reduction.
[0098] Step 4: Determine multiple reward values based on the abandoned wind and photovoltaic power consumed by the energy conversion subsystem and the methane conversion efficiency, and sum up the weighted values of each reward value to obtain the reward for the energy conversion equipment.
[0099] Among them, by rewarding the energy conversion equipment for consuming abandoned wind and photovoltaic power, it promotes the full utilization of clean energy and reduces energy waste. At the same time, it rewards the efficiency of the equipment in converting into methane, encourages improving the energy conversion efficiency, and enables more electric energy to be effectively converted into natural gas that can be stored and utilized. This helps to enhance the energy storage capacity of the system, improve the comprehensive energy utilization efficiency, and enhance the system's ability to cope with energy supply and demand fluctuations.
[0100] Step 5: Determine the reward value and penalty value respectively based on the actual reduced electricity load of the user and the incentive expenditure, and sum up the weighted values of the reward value and the penalty value to obtain the user demand response reward.
[0101] Specifically, it rewards the user for actually reducing the electricity load, directly encouraging the user to actively respond to the system demand, participate in peak shaving and valley filling, and relieve the pressure on the power grid during peak hours. At the same time, it penalizes the incentive expenditure to avoid excessive incentives leading to too high costs, and prompts the system to find the most cost-effective incentive strategy during the user demand response process, improving the sustainability and effectiveness of user participation in demand response.
[0102] It can be understood that the reward functions of each part are interrelated and jointly act on each link of the integrated energy system. Through the rewards and penalties for different equipment and user behaviors, it guides the system to optimize the operation strategy from multiple dimensions such as power generation, energy conversion, and load regulation. Furthermore, it helps the model optimize the parameters of the continuous strategy network, and finally obtains an optimized operation strategy that can meet the generation of multiple dimensions to be optimal.
[0103] S103: Through the reinforcement learning algorithm, based on the preprocessed system model, perform low-carbon dispatching optimization.
[0104] Specifically, this step includes the following sub-steps:
[0105] Step 1. Read the status information related to the tasks of each subsystem through the agents of each subsystem;
[0106] It can be understood that each subsystem agent first reads the relevant status information. Based on this information, scheduling actions can be generated through the policy network, realizing the key step from system perception to decision-making output.
[0107] For example, the wind-solar power generation agent determines the power generation allocation plan according to status information such as wind-solar power generation prediction values and real-time grid load demands; the thermal power unit agent determines the adjustment strategy of the output power and carbon capture rate based on its own power output data, carbon capture equipment operation data, and grid load conditions.
[0108] Step 2. Generate scheduling actions through the policy network according to the status information, and feedback the scheduling actions to the system model;
[0109] Step 3. Indicate each subsystem to respond to the scheduling actions respectively. After the scheduling actions are executed, calculate the reward value according to the execution situation through the value network, using the total system operation cost, carbon emissions, and wind-solar power consumption rate as the measurement indicators;
[0110] Specifically, after the generated scheduling actions are feedback to the system model, each subsystem executes the corresponding actions, changing the system operation state. Subsequently, the value network calculates the reward value according to the execution situation, using the total system operation cost, carbon emissions, and wind-solar power consumption rate as the measurement indicators. This process quantifies the actual operation effect of the system, providing a basis for subsequent optimization.
[0111] For example, if the wind-solar power generation agent supplies more wind-solar power to the grid, reducing the system operation cost and increasing the wind-solar power consumption rate, it will receive a higher reward; if the thermal power unit agent effectively supplements the grid load gap and increases the carbon dioxide capture volume, it will also receive corresponding rewards.
[0112] Step 4. Obtain historical data from the experience replay pool, and continuously optimize the parameters of the policy network according to the historical data and the reward value. Generate scheduling actions for low-carbon scheduling through the continuously optimized system model.
[0113] Among them, obtain historical data from the experience replay pool, and continuously optimize the parameters of the policy network in combination with the current reward value. Enable the agent to learn from past experience and continuously improve the decision-making strategy. For example, if a certain scheduling action performs well (receives a high reward) in the historical data, it will be more inclined to select a similar action in subsequent decisions, and vice versa. Through continuous iterative optimization, the scheduling actions generated by the system model will be more and more conducive to achieving the low-carbon scheduling goal and improving the overall performance of the system.
[0114] In an exemplary embodiment, the update process of the policy network parameters further includes the following sub-steps:
[0115] Step4.1, packetize and store the historical feedback data accumulated by the integrated energy system when performing feedback actions to obtain an experience replay pool. The historical feedback data includes the current state of the agent, the actions performed, the rewards obtained, the new state after performing the actions, and the parameters of the historical policy network;
[0116] Among them, packetizing and storing the historical feedback data accumulated by the integrated energy system when performing feedback actions to form an experience replay pool provides a rich data basis for subsequent learning. Information such as the agent state, actions performed, rewards obtained, new state, and historical policy network parameters included comprehensively records the past operation and decision-making of the system.
[0117] Step4.2, randomly extract target data from the experience replay pool. According to the target data, obtain the probability value when the new policy network and the old policy network select the same action, and use the ratio of the probability values as the calculation weight for the next update of the policy network;
[0118] Randomly extract target data from the experience replay pool and obtain the ratio of the probability values of the new and old policy networks selecting the same action as the calculation weight. The significance of this step is to measure the improvement or difference of the new policy network by comparing the decision-making tendencies of the new and old policy networks in the same situation. If the new policy network is more inclined to select actions that lead to high rewards (higher probability values), the calculation weight will make the new policy network adjust more in the favorable direction when updating parameters. Conversely, it will suppress the potentially unfavorable update direction, thus guiding the policy network to develop towards a better decision-making direction.
[0119] Step4.3, subtract the reward value of the next state from the overall reward value of the current state to obtain the calculation step size for the next update of the policy network;
[0120] Among them, the update of the policy network is based on the combined action of global rewards and local rewards. Global rewards provide feedback on the overall performance of the system, guiding the policy network to adjust parameters in the direction of optimizing the overall goal (such as low carbon, economy, supply-demand balance, etc.). Local rewards evaluate the specific behaviors and performances of each subsystem, enabling the policy network to refine the decision optimization direction according to the characteristics and tasks of each subsystem.
[0121] Subtract the next state reward value from the current state reward value to obtain the calculation step, which reflects the change in the reward after the system executes an action. If the reward value increases, the step is positive, indicating that the current action has a positive impact on the system, and the policy network should be further optimized in this direction; if the reward value decreases, the step is negative, suggesting that the decision-making direction may need to be adjusted. The calculation step provides a basis for adjusting the amplitude of the policy network parameter update, enabling the update process to be reasonably adjusted according to the actual reward change.
[0122] Step4.4, based on the calculated weight and the calculation step, use the gradient descent method to update the parameters of the policy network to obtain a new policy network.
[0123] Based on the calculated weight and the calculation step, use the gradient descent method to update the policy network parameters to obtain a new policy network. The gradient descent method gradually adjusts the policy network parameters along the direction that maximizes the reward function according to the calculated weight and the step. The calculated weight determines the direction and intensity of the update, and the calculation step determines the step size of each update. By continuously repeating this process, the policy network can gradually learn the optimal decision-making strategy to adapt to the complex operating environment of the integrated energy system and achieve better low-carbon scheduling effects.
[0124] It can be understood that the entire process of updating the policy network parameters realizes the intelligent learning ability of the agent. By reviewing and analyzing historical experience (experience replay pool) and combining the current system operation feedback (reward value change), the policy network parameters are continuously adjusted, enabling the agent to automatically adapt to the dynamic changes of the system and optimize the decision-making process. In an optional embodiment, the output of the optimal scheduling method includes: the operation scheduling strategies of each device, as well as the real-time updated total carbon emissions and operating costs of the system.
[0125] Through the above steps S101 to S103, the dynamic optimization and real-time decision-making of the integrated energy system are realized. The scheduling actions are adjusted in a timely manner according to the continuously changing state information to adapt to various working conditions (such as fluctuations in wind and solar power generation, changes in load demand, etc.). In each decision-making cycle, through the loop of generating actions, executing actions, calculating rewards, and optimizing parameters by the policy network, the system always operates in the direction of low-carbon, economic, and efficient, ensuring optimal scheduling under different operating conditions.
[0126] In addition, the continuously optimized system model helps to improve energy utilization efficiency, reduce carbon emissions, and enhance system stability, thereby promoting the sustainable development of the integrated energy system. By encouraging each subsystem to pay attention to environmental protection and economic indicators while meeting energy demands, the balance between energy supply, environmental protection, and economic benefits is achieved. This not only meets the current requirements of the "dual carbon" goal but also provides a feasible technical path and practical experience for the intelligent and low-carbon development of future energy systems.
[0127] Through the method of the present invention, the integrated energy system can dynamically adapt to environmental changes (such as fluctuations in wind and solar power generation, changes in grid load demand, user behavior response, etc.), continuously optimize the dispatching strategies of each subsystem, effectively improve the utilization rate of wind and solar power generation, reduce the phenomenon of wind and light abandonment, optimize the power generation and carbon capture efficiency of thermal power units, maximize the utilization of excess wind and solar power by P2G equipment, and at the same time reduce peak loads through user demand response, so as to achieve low-carbon and efficient operation of the integrated energy system as a whole.
[0128] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0129] The embodiment of the present application also provides a low-carbon dispatching system for an integrated energy system based on carbon capture and demand response. Figure 4 is a structural block diagram of a low-carbon dispatching system for an integrated energy system based on carbon capture and demand response according to an embodiment of the present application, as Figure 4 shown, the system includes: a model establishment module 40, a preprocessing module 41, and a dispatching optimization module 42, where:
[0130] The model establishment module 40 is used to establish a system model of the integrated energy system, where the integrated energy system includes: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem, and a user demand response subsystem, and among them, the thermal power subsystem is configured with carbon capture equipment;
[0131] The preprocessing module 41 is used to preprocess the system model, including: defining the state space, action space, and reward function corresponding to the reinforcement learning algorithm for each subsystem of the integrated energy system, and dividing each subsystem into independent optimization units through multi-agent partitioning;
[0132] The dispatching optimization module 42 is used to perform low-carbon dispatching optimization based on the preprocessed system model through the reinforcement learning algorithm.
[0133] In addition, Figure 5 is a diagram of the combined operation mode of a low-carbon integrated energy system according to an embodiment of the present application, as Figure 5As shown, through this system, the dynamic optimization and real-time decision-making of the integrated energy system are realized. According to the continuously changing state information, the dispatching actions are adjusted in a timely manner to adapt to various working conditions (such as fluctuations in wind and solar power generation, changes in load demand, etc.). In each decision-making cycle, through the loop of generating actions, executing actions, calculating rewards, and optimizing parameters by the policy network, the system always operates in the direction of low-carbon, economic, and efficient, ensuring optimal dispatching under different operating conditions. It effectively improves the utilization rate of wind and solar power generation, reduces the phenomenon of wind and light abandonment, optimizes the power generation and carbon capture efficiency of thermal power units, maximizes the utilization of excess wind and solar power by P2G equipment, and at the same time reduces peak loads through user demand response, realizing the low-carbon and efficient operation of the integrated energy system as a whole.
[0134] In one embodiment, Figure 6 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. As Figure 6 shown, an electronic device is provided. The electronic device can be a server. The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store data. The network interface of the electronic device is used to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a low-carbon dispatching method for an integrated energy system based on carbon capture and demand response.
[0135] Those skilled in the art can understand that Figure 6 the structure, which is only a block diagram of some structures related to the solution of the present application, does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0136] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0137] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0138] In addition, in combination with the above-mentioned method for centralized power prediction of a wind farm based on big data, the embodiments of the present application can provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the above-mentioned methods for low-carbon scheduling of an integrated energy system based on carbon capture and demand response is implemented.
[0139] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0140] The above embodiments only illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A low-carbon dispatching method for an integrated energy system based on carbon capture and demand response, characterized in that: Low-carbon dispatch optimization of an integrated energy system is performed by using a reinforcement learning algorithm of a strategy-value network based on experience replay, the method comprising: Establishing a system model of an integrated energy system, wherein the integrated energy system comprises: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem and a user demand response subsystem, wherein the thermal power subsystem is equipped with a carbon capture device; Preprocessing the system model includes: defining the state space, action space and reward function corresponding to the reinforcement learning algorithm for each subsystem of the integrated energy system, and dividing each subsystem into independent optimization units through multi-agent partitioning; Through the reinforcement learning algorithm, low-carbon scheduling optimization is performed based on the system model after the preprocessing.
2. The method according to claim 1, characterized in that: Based on the system model after the preprocessing, low-carbon scheduling optimization includes: Through the intelligent agents of each subsystem, read the status information related to the tasks of each subsystem; Generate a scheduling action through a policy network according to the state information, and feed the scheduling action back to the system model; Instructing each subsystem to respond to the dispatching action respectively, and after the dispatching action is completed, calculating the reward value through the value network according to the execution status, taking the total operating cost of the system, carbon emissions and wind and solar power consumption rate as measurement indicators; Historical data is obtained from the experience replay pool, and the parameters of the policy network are continuously optimized according to the historical data and the reward value, and a scheduling plan for low-carbon scheduling is generated through a continuously optimized system model.
3. The method according to claim 1, characterized in that The state space includes the global state of the system and the local state of each subsystem, where: The global state of the system includes: wind and solar power generation forecast value, real-time load demand of the power grid, total carbon emissions and carbon trading price; The local status of each subsystem includes: wind power generation data, photovoltaic power generation data, wind abandonment data and solar abandonment data, thermal power unit power output data, carbon capture equipment operation data, energy conversion subsystem operating power and gas production, and user demand response.
4. The method according to claim 1, characterized in that: Defining the action space for each subsystem includes: For the intelligent bodies of the wind subsystem and the photovoltaic subsystem: formulate a power generation priority allocation plan, and when the grid demand load is greater than or equal to the wind and solar power generation, all the wind and solar power generation is used for grid load; When the grid load demand is less than the wind and solar power generation, the wind and solar power generation is used to meet the grid load demand, and the power that does not meet the grid load demand is used to drive the energy conversion subsystem; For the intelligent body of the thermal power subsystem: by adjusting its output power and carbon capture rate, when the wind and solar power generation cannot meet the load demand of the power grid, the thermal power subsystem is used to supplement the gap of the power grid load demand; When the power generation of the thermal power unit can meet the shortfall of the load demanded by the power grid and there is surplus power, the carbon capture device is driven by the surplus power to capture the carbon dioxide in the flue gas emitted by the thermal power generation system; For the intelligent agent of the energy conversion subsystem: when the wind and solar power generation is in excess, the energy conversion subsystem uses the excess electricity to produce hydrogen, and the hydrogen reacts with the carbon dioxide in the carbon capture device to generate methane gas; For the intelligent agent of the user demand response subsystem: through preset incentive measures, guide users to adjust their electricity consumption behavior to reduce the electricity load during peak hours or shift the electricity demand to off-peak hours.
5. The method according to claim 4, characterized in that The action space of the agent of the user demand response subsystem includes: incentive intensity dimension, user type dimension and response period dimension, wherein: The incentive intensity dimension is used to determine the price discount or reward intensity provided to users, including low incentive level, medium incentive level and high incentive level; The user type dimension is used to classify users into residential users, industrial users and commercial users based on the characteristics of different users, and to set corresponding action strategies according to different user types; The response period dimension is used to select a response period for the incentive effect according to the operating status of the power grid, and the response period includes a peak period, a valley period and a load balancing period.
6. The method according to claim 1, characterized in that Defining the reward function includes: A global reward is obtained by weighted summation based on carbon emissions, operating costs, wind and solar power consumption rates, and load balance parameters; Determine reward values and penalty values according to the online proportions and wind abandonment proportions or solar abandonment proportions of the wind subsystem and the photovoltaic subsystem, respectively, and obtain wind power generation rewards and photovoltaic power generation rewards by weighted summing the reward values and penalty values; Determine a plurality of thermal power reward and penalty values respectively according to the load supplement amount, carbon capture amount, and fuel usage of the thermal power unit, and weighted sum each thermal power reward and penalty value to obtain a reward for the thermal power unit; Determine multiple conversion reward values according to the abandoned wind and solar power and methane conversion efficiency absorbed by the energy conversion subsystem, and obtain the energy conversion equipment reward by weighted summing up the conversion reward values; According to the actual reduction in power load and incentive expenditure measured by the user, the user demand reward value and the user demand penalty value are determined respectively, and the user demand response reward is obtained by weighted summing the user demand reward value and the user demand penalty value.
7. The method according to claim 1, characterized in that The updating process of the policy network parameters includes: The historical feedback data obtained by the integrated energy system executing the feedback action is packaged and stored to generate an experience replay pool, wherein the historical feedback data includes the current state of the agent, the action executed, the reward obtained, the new state after the action is executed, and the parameters of the historical strategy network; The target data is randomly extracted from the experience replay pool, and based on the target data, a probability value when the new policy network and the old policy network select the same action is obtained, and the ratio of the probability values is used as a calculation weight for updating the policy network in the next step; Subtract the total reward value of the next state from the total reward value of the current state to obtain the calculation step size for updating the policy network in the next step; Based on the calculation weight and the calculation step size, the parameters of the policy network are updated using a gradient descent method to obtain a new policy network.
8. A low-carbon dispatching system for an integrated energy system based on carbon capture and demand response, characterized in that: The low-carbon dispatch optimization of the integrated energy system is performed by a reinforcement learning algorithm of a strategy-value network based on experience replay. The system includes: a model building module, a preprocessing module and a dispatch optimization module, wherein: The model building module is used to build a system model of an integrated energy system, wherein the integrated energy system includes: a thermal power subsystem, a wind power subsystem, a photovoltaic subsystem, an energy conversion subsystem and a user demand response subsystem, wherein the thermal power subsystem is equipped with a carbon capture device; The preprocessing module is used to preprocess the system model, including: defining the state space, action space and reward function corresponding to the reinforcement learning algorithm for each subsystem of the integrated energy system, and dividing each subsystem into independent optimization units through multi-agent division; The scheduling optimization module is used to perform low-carbon scheduling optimization based on the pre-processed system model through the reinforcement learning algorithm.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Flue gas desulfurization system coordination control method based on multi-agent reinforcement learning
CN121550817A