Energy scheduling method and device in park, storage medium and electronic equipment
Through the target strategy network trained by deep deterministic strategy gradient algorithm, the global optimal solution problem of energy scheduling of multiple subjects in the park is solved, the dual optimization of the economics of the park's energy system and carbon reduction goals is achieved, and the intelligence level of energy management is improved.
Patent Information
- Application Number
- CN202510568837.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The energy scheduling objective functions and constraints of multiple subjects in the park affect each other, making it difficult to find the global optimal solution for the optimization problem, resulting in poor energy saving and carbon reduction effects.
The target policy network is trained using the deep deterministic policy gradient algorithm (DDPG), and iteratively updates the initial policy network with multiple constraints and sample data, and outputs action variables for energy scheduling, including the action decisions of energy supply, conversion and storage systems.
The dual optimization of the economics and carbon reduction goals of the park's energy system has been achieved, and it can respond to changes in energy demand in real time, adapt to market and weather uncertainties, reduce operating costs and carbon emissions, and improve the level of intelligent energy management.
Smart Images

Figure CN120471373A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to an energy scheduling method, device, storage medium, and electronic equipment within a park. Background Art
[0002] As a typical integrated energy system, the Park Integrated Energy System (PIES) aggregates multiple energy load demands, including electricity, heat, and gas, and has significant energy-saving and carbon-reduction effects, and helps promote the efficient consumption of new energy. However, park energy systems face multiple challenges, such as complex coupling relationships between devices, uncertain operating characteristics, and low load forecasting accuracy. As a result, traditional decision-making methods, such as those based on game theory and robust optimization, often lack adaptability and portability in multiple scenarios. At the same time, faced with the coordination and rolling optimization of multiple agents and multiple time scales within the park, existing single-stage optimization, two-level optimization, and stochastic optimization methods cannot effectively solve the complex non-convex problems in multi-agent evolutionary games and the random fluctuations of dynamic response variables, making it difficult to achieve global optimal results.
[0003] In related technologies, due to the existence of multiple entities within the park, the objective functions and constraints corresponding to the energy scheduling of each entity influence each other, making it difficult to find a global optimal solution for the optimization problem, resulting in poor energy conservation and carbon reduction effects in the park. No effective solution has been proposed yet. Summary of the Invention
[0004] The main purpose of this application is to provide an energy scheduling method, device, storage medium and electronic equipment within a park, so as to solve the problem in related technologies that due to the existence of multiple entities in the park, the objective functions and constraints corresponding to the energy scheduling of each entity influence each other, making it difficult to find a global optimal solution for the optimization problem, resulting in poor energy conservation and carbon reduction effects in the park.
[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for energy scheduling within a park is provided, which includes: collecting the operating status of multiple energy systems in the park within a target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system, and the energy storage system includes at least: power storage equipment, gas storage equipment, and heat storage equipment; based on a deep deterministic policy gradient algorithm, the state variables are input into a target policy network, and action variables are output, wherein the action variables are used to characterize the action decision information of the park scheduling the multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network; and energy scheduling is performed on the multiple energy systems in the park based on the action variables.
[0006] Furthermore, the target policy network is obtained by iterative training of the following steps: constructing an objective function based on the operating costs of the multiple energy systems and the total carbon emissions of the park; constructing a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; constructing the target value network based on the reward function, and constructing an initial policy network based on the target value network; using the deep deterministic policy gradient algorithm, the multiple constraints and the multiple sample data to iteratively update the initial policy network to obtain the target policy network.
[0007] Furthermore, the deep deterministic policy gradient algorithm, the multiple constraints and the multiple sample data are used to iteratively update the initial policy network to obtain the target policy network, including: constructing an electric energy constraint based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment and the electric load stored by the power storage equipment; wherein the first type of energy conversion equipment is a device that converts gas into electric energy; constructing a thermal energy constraint based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment and the heat load stored by the heat storage equipment; wherein the second type of energy conversion equipment is a device that converts gas into thermal energy; and iteratively updating the initial policy network based on the electric energy constraint, the thermal energy constraint, the constraint of the power storage equipment, the output power constraint of the energy supply system and the output power constraint of the energy storage system to obtain the target policy network.
[0008] Furthermore, the constraints of the energy storage device are obtained by the following steps: constructing the state of charge constraint of the energy storage device based on the energy storage state of charge of the energy storage device within a preset time period, the energy storage charging and discharging efficiency of the energy storage device, and the maximum energy storage capacity of the energy storage device; constructing the usage time constraint of the energy storage device based on the charging power of the energy storage device within a preset time period, the discharging power of the energy storage device, the usage time limit of the energy storage device, and the maximum number of charge and discharge cycles of the energy storage device within the usage time limit; determining the constraint of the energy storage device based on the usage time constraint of the energy storage device and the state of charge constraint of the energy storage device.
[0009] Furthermore, the target value network is constructed based on the reward function, and the initial policy network is constructed based on the target value network, including: determining a first loss function based on the reward function, the state variables, the action variables corresponding to the state variables, and the expected return information; calculating the network parameters of the target value network during iterative update by performing a gradient operation on the first loss function, and constructing the target value network; determining the second loss function of the initial policy network based on the maximum value of the target value network; calculating the network parameters of the initial policy network during iterative update by performing a gradient operation on the second loss function, and constructing the initial policy network.
[0010] Furthermore, an objective function is constructed based on the operating costs of the multiple energy systems and the total carbon emissions of the park, including: determining the gas purchase cost of the park based on the gas purchase price of the park, the output power of the first type of energy conversion equipment, and the conversion efficiency of the third type of energy conversion equipment; wherein the third type of energy conversion equipment represents equipment that converts gas into electrical energy or thermal energy; determining the energy storage equipment in the energy storage system, and determining the operating cost of the energy storage system based on the number of the energy storage equipment and the unit operating cost of the energy storage equipment; determining the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; constructing the objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid, and the carbon emission coefficient of the first type of energy conversion equipment.
[0011] Furthermore, the reward function includes at least: an instant reward function and a cumulative reward function; a reward function is constructed based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data, including: constructing the state variables based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge status information of the energy storage device, the output power prediction value of the wind power system, and the output power prediction value of the photovoltaic system; constructing the action variables based on the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; constructing the reward function of the energy scheduling of the park at any time based on the state variables, the action variables and the objective function to obtain the instant reward function; constructing the cumulative reward function based on the instant reward function and preset parameters of the energy scheduling of the park within the preset time period.
[0012] In order to achieve the above-mentioned purpose, according to another aspect of the present application, an energy scheduling device within a park is provided, which includes: a collection unit, which is used to collect the operating status of multiple energy systems in the park within a target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system, and the energy storage system includes at least: power storage equipment, gas storage equipment, and heat storage equipment; a calculation unit, which is used to input the state variables into a target policy network based on a deep deterministic policy gradient algorithm and output action variables, wherein the action variables are used to characterize the action decision information of the park scheduling the multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network; and a scheduling unit, which is used to perform energy scheduling on the multiple energy systems in the park based on the action variables.
[0013] Furthermore, the device also includes: a first construction unit, used to construct an objective function based on the operating costs of the multiple energy systems and the total carbon emissions of the park; a second construction unit, used to construct a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; a third construction unit, used to construct the target value network based on the reward function, and to construct an initial policy network based on the target value network; an update unit, used to iteratively update the initial policy network using the deep deterministic policy gradient algorithm, the multiple constraints and the multiple sample data to obtain the target policy network.
[0014] Furthermore, the update unit includes: a first construction subunit, used to construct an electric energy constraint condition based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment and the electric load stored in the power storage equipment; wherein, the first type of energy conversion equipment is a device that converts gas into electric energy; a second construction subunit, used to construct a thermal energy constraint condition based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment and the heat load stored in the heat storage equipment; wherein, the second type of energy conversion equipment is a device that converts gas into thermal energy; an update subunit, used to iteratively update the initial strategy network based on the electric energy constraint condition, the thermal energy constraint condition, the constraint condition of the power storage equipment, the output power constraint condition of the energy supply system and the output power constraint condition of the energy storage system to obtain the target strategy network.
[0015] Furthermore, the update unit also includes: a third construction subunit, used to construct the state of charge constraint condition of the power storage device based on the energy storage state of charge of the power storage device within a preset time period, the energy storage charging and discharging efficiency of the power storage device, and the maximum energy storage capacity of the power storage device; a fourth construction subunit, used to construct the usage time constraint condition of the power storage device based on the charging power of the power storage device within a preset time period, the discharging power of the power storage device, the usage time limit of the power storage device, and the maximum number of charge and discharge cycles of the power storage device within the usage time limit; a first determination subunit, used to determine the constraint condition of the power storage device based on the usage time constraint condition of the power storage device and the state of charge constraint of the power storage device.
[0016] Furthermore, the third construction unit includes: a second determination subunit, used to determine a first loss function based on the reward function, the state variable, the action variable corresponding to the state variable and the expected return information; a first calculation subunit, used to calculate the network parameters of the target value network during iterative update by performing a gradient operation on the first loss function, and construct the target value network; a third determination subunit, used to determine the second loss function of the initial policy network based on the maximum value of the target value network; a second calculation subunit, used to calculate the network parameters of the initial policy network during iterative update by performing a gradient operation on the second loss function, and construct the initial policy network.
[0017] Furthermore, the first construction unit includes: a fourth determination subunit, used to determine the gas purchase cost of the park based on the gas purchase price of the park, the output power of the third type of energy conversion equipment, and the conversion efficiency of the first type of energy conversion equipment; wherein the third type of energy conversion equipment refers to equipment that converts gas into electrical energy or thermal energy; a fifth determination subunit, used to determine the energy storage equipment in the energy storage system, and determine the operating cost of the energy storage system based on the number of the energy storage equipment and the unit operating cost of the energy storage equipment; a sixth determination subunit, used to determine the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; a fifth construction subunit, used to construct the objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid, and the carbon emission coefficient of the first type of energy conversion equipment.
[0018] Furthermore, the reward function includes at least: an instant reward function and a cumulative reward function; the second construction unit includes: a sixth construction subunit, which is used to construct the state variable based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge state information of the energy storage device, the output power prediction value of the wind system, and the output power prediction value of the photovoltaic system; a seventh construction subunit, which is used to construct the action variable based on the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; an eighth construction subunit, which is used to construct the reward function of the energy scheduling of the park at any time based on the state variable, the action variable and the objective function to obtain the instant reward function; a ninth construction subunit, which is used to construct the cumulative reward function based on the instant reward function and preset parameters of the energy scheduling of the park within the preset time period.
[0019] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a computer program product is provided, including a computer program, which, when executed by a processor, implements any of the above-mentioned energy scheduling methods within the park, and which, when executed by a processor, implements the steps of the energy scheduling method within the park described in each embodiment of the present application.
[0020] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes stored computer instructions, wherein when the computer instructions are executed by a processor, any one of the above-mentioned energy scheduling methods within the park is implemented.
[0021] In order to achieve the above-mentioned purpose, according to one aspect of the present application, an electronic device is provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement any one of the above-mentioned energy scheduling methods within the park.
[0022] The present application adopts the following steps: collecting the operating status of multiple energy systems in the park during a target time period to obtain state variables, wherein the multiple energy systems include at least an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least an external power grid, a wind power system, and a photovoltaic system, and the energy storage system includes at least an electrical storage device, a gas storage device, and a heat storage device; inputting the state variables into a target policy network based on a deep deterministic policy gradient algorithm, and outputting action variables, wherein the action variables are used to characterize the action decision information of the park scheduling the multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network; and energy scheduling of the multiple energy systems in the park is performed based on the action variables, thereby solving the problem in the related art that due to the existence of multiple subjects in the park, the objective function corresponding to the energy scheduling of each subject and the constraints interact with each other, making it difficult to find a global optimal solution for the optimization problem, resulting in poor energy conservation and carbon reduction effects in the park.
[0023] By monitoring and recording the power, wind, and photovoltaic supply within the park, as well as the dynamics of electricity, gas, and heat storage, a data set reflecting the overall status of the energy system is established, enabling the intelligent agent to make more accurate decisions based on current and predicted energy demand and supply. The target policy network is trained through the DDPG algorithm, which can process continuous action space and effectively respond to the dynamic balance and complex coupling relationships within the system. Even in the face of unexpected load fluctuations or changes in energy prices, the system can maintain efficient operation. By setting multiple constraints, it is ensured that the action variables output by the intelligent agent can not only reduce costs and emissions, but also comply with physical and operational limitations, avoiding unrealistic or harmful scheduling decisions. Through interaction with the target value network, the intelligent agent continuously learns and optimizes its strategy to obtain the highest return (that is, the above-mentioned return information), that is, the minimum total cost and carbon emissions.
[0024] In summary, the agent's decision-making has achieved dual optimization of the park's energy system's economic efficiency and carbon reduction goals. It not only responds to changes in the park's energy demand in real time, but also predicts and adapts to market and weather uncertainties, reducing both operating costs and carbon emissions. The solution demonstrates its flexibility and robustness across diverse scenarios, providing the park with a stable, economical, and environmentally friendly energy supply and significantly enhancing the intelligent level of energy management. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0026] Figure 1 This is a flow chart of the energy scheduling method within the park provided in Example 1 of the present application;
[0027] Figure 2 This is a flowchart of an optional park economic carbon reduction optimization method based on a deep reinforcement learning algorithm provided in Example 1 of the present application;
[0028] Figure 3 Schematic diagram of an optional park integrated energy optimization and scheduling system provided in accordance with the first embodiment of the present application;
[0029] Figure 4 This is an optional campus structure diagram based on the Markov Decision Process (MDP) provided in accordance with the first embodiment of the present application;
[0030] Figure 5 This is a schematic diagram of an energy scheduling device within a park provided according to Example 2 of the present application;
[0031] Figure 6 This is a schematic diagram of energy scheduling electronic equipment within a park provided according to Example 5 of the present application. DETAILED DESCRIPTION
[0032] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0033] It should be noted that the user information (including but not limited to user device information, user personal information, collected data, used data, generated data, processed data, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, collected information, used information, generated information, processed information, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or institution through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or institution.
[0034] It should be noted that this application provides users with corresponding operation entrances for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.
[0035] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0037] Example 1
[0038] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 This is a flow chart of the energy scheduling method within the park provided in Example 1 of the present application. Figure 1 As shown, the method includes the following steps:
[0039] Step S101 collects the operating status of multiple energy systems in the park during the target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system. The energy supply system includes at least: an external power grid, a wind power system, and a photovoltaic system. The energy storage system includes at least: electricity storage equipment, gas storage equipment, and heat storage equipment.
[0040] In this first embodiment, the operational information of each energy system within the park during the current time period (i.e., the aforementioned target time period) must first be collected to generate comprehensive state variables. These state variables cover the real-time status of the energy supply system, such as the power supply capacity of the external power grid and the power production of wind turbines and photovoltaic panels. They also include the operating status of the energy conversion system, including the conversion efficiency and output power of power-to-gas (P2G), gas turbines (GT), gas boilers (GB), and electric boilers (EB). Furthermore, attention must be paid to the charge and discharge status and energy storage level of the power, gas, and heat storage devices in the energy storage system.
[0041] The execution subject of the energy scheduling method within the park provided in the first embodiment can be an agent based on a deep reinforcement learning algorithm, specifically, an intelligent optimization system that uses a deep deterministic policy gradient (DDPG) algorithm. The agent is responsible for receiving the status information of the energy system within the park, and learning and outputting the optimal energy scheduling strategy through a deep reinforcement learning framework, especially iterative updates of the value network and the policy network. The agent's decision directly affects the operation of each energy device in the park, thereby achieving the dual optimization goals of economic cost and carbon emissions.
[0042] In step S102, based on the deep deterministic policy gradient algorithm, the state variables are input into the target policy network and the action variables are output, where the action variables are used to characterize the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and the target value network; the initial policy network is responsible for outputting the action decision information of the park's energy scheduling under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network.
[0043] In this first embodiment, the Deep Deterministic Policy Gradient (DDPG) algorithm uses state variables as input to drive the target policy network to generate action variables. The action variables reflect the specific operational decisions when scheduling various energy systems in the park, such as adjusting the power of energy storage devices, gas turbines, and gas boilers.
[0044] The target policy network is an optimization model that uses a large amount of sample data, using future reward predictions from the target value network and iteratively updating the initial policy network based on multiple constraints. The policy network outputs the corresponding energy scheduling decision for any given state variable. The target value network evaluates the long-term benefits of these energy scheduling decisions and guides the policy network to maximize the cumulative return (i.e., the aforementioned return information), thereby minimizing the park's overall costs and carbon emissions.
[0045] Step S103 : performing energy scheduling on multiple energy systems in the park according to the action variables.
[0046] In this first embodiment, precise scheduling instructions are executed for the energy system within the park based on the action variables output by the deep reinforcement learning algorithm. These action variables can include operational recommendations for each energy device, such as adjusting the charge and discharge power of energy storage devices, changing the production capacity of gas turbines or boilers, and controlling the operating status of power-to-gas equipment. Based on this decision information, electricity, heat, and gas resources are automatically allocated, ensuring supply and demand matching while reducing operating costs and carbon emissions.
[0047] Specifically, after the intelligent agent calculates and determines the action variables based on the current state variables, it will command the energy supply, conversion and storage systems, dynamically adjust their operating modes to adapt to changes in load demand, and optimize energy utilization efficiency to achieve scheduling goals that are both economical and environmentally friendly.
[0048] In summary, the energy scheduling method within the park provided in Example 1 of the present application obtains state variables by collecting the operating status of multiple energy systems in the park within a target time period, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system, and the energy storage system includes at least: power storage equipment, gas storage equipment, and heat storage equipment; based on the deep deterministic policy gradient algorithm, the state variables are input into the target policy network, and action variables are output, wherein the action variables are used to characterize the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and the target value network; the initial policy network is responsible for outputting the action decision information of the park's energy scheduling under any state variable; the target value network is used to predict the return information of the action variables output by the policy network; energy scheduling is performed on multiple energy systems in the park based on the action variables, which solves the problem in the related technology that due to the existence of multiple subjects in the park, the objective function corresponding to the energy scheduling of each subject and the constraints affect each other, making it difficult to find the global optimal solution for the optimization problem, resulting in poor energy saving and carbon reduction effects in the park.
[0049] By monitoring and recording the power, wind, and photovoltaic supply within the park, as well as the dynamics of electricity, gas, and heat storage, a data set reflecting the overall status of the energy system is established, enabling the intelligent agent to make more accurate decisions based on current and predicted energy demand and supply. The target policy network is trained through the DDPG algorithm, which can process continuous action space and effectively respond to the dynamic balance and complex coupling relationships within the system. Even in the face of unexpected load fluctuations or changes in energy prices, the system can maintain efficient operation. By setting multiple constraints, it is ensured that the action variables output by the intelligent agent can not only reduce costs and emissions, but also comply with physical and operational limitations, avoiding unrealistic or harmful scheduling decisions. Through interaction with the target value network, the intelligent agent continuously learns and optimizes its strategy to obtain the highest return (that is, the above-mentioned return information), that is, the minimum total cost and carbon emissions.
[0050] In summary, the agent's decision-making has achieved dual optimization of the park's energy system's economic efficiency and carbon reduction goals. It not only responds to changes in the park's energy demand in real time, but also predicts and adapts to market and weather uncertainties, reducing both operating costs and carbon emissions. The solution demonstrates its flexibility and robustness across diverse scenarios, providing the park with a stable, economical, and environmentally friendly energy supply and significantly enhancing the intelligent level of energy management.
[0051] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, the target policy network is obtained by iterative training through the following steps: constructing an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park; constructing a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; constructing a target value network based on the reward function, and constructing an initial policy network based on the target value network; and iteratively updating the initial policy network using a deep deterministic policy gradient algorithm, multiple constraints, and multiple sample data to obtain a target policy network.
[0052] In this first embodiment, the deep deterministic policy gradient (DDPG) algorithm is used to construct and optimize the objective function, reward function, value network and policy network to achieve the goal of minimizing the cost and carbon emissions of the campus energy system.
[0053] First, an objective function is designed based on the park's energy system's operating costs and total carbon emissions, aiming to minimize the park's total energy expenditure (including the procurement and conversion costs of electricity, heat, and gas) and the environmental cost (carbon emissions). At the same time, the impact of costs and emissions can be adjusted by setting weighting factors u1 and u2 to ensure decision-making flexibility.
[0054] Next, a reward function is constructed based on the objective function, the state variables, and the action variables of multiple sample data collected within a preset time period. The reward function can be designed to be the negative value of the objective function, meaning that the lower the cost and the lower the carbon emissions, the higher the reward, thereby guiding the agent to learn a more optimal scheduling strategy. On this basis, a target value network is established, which is used to predict the reward value in the future state after taking a specific action variable. The output value Q(s, a) of the value network evaluates the expected return of performing action a in state s, providing feedback to the policy network. The initial policy network is responsible for outputting an action decision based on the current state. It gradually learns and adjusts the decision-making strategy through interaction with the value network.
[0055] Finally, the initial policy network is iteratively updated using a large amount of sample data through a deep deterministic policy gradient algorithm. In each round of training, the agent outputs actions based on the current state and receives immediate feedback (rewards) through interaction with the environment (i.e., the actual operation of the campus energy system). These experience data (state, action, reward, next state) are then stored in a preset storage space. The algorithm randomly extracts samples from the preset storage space to update the policy network and value network, and gradually approaches the optimal strategy by minimizing the loss function of the value network and maximizing the immediate and future rewards of the policy network. This process is repeated until the policy network can stably output scheduling decisions that minimize the objective function value, forming a target policy network and providing an intelligent solution for campus energy optimization scheduling.
[0056] Through the above steps, a campus energy optimization and scheduling framework based on deep reinforcement learning can be constructed. It achieves the comprehensive optimization of economic costs and carbon emissions by continuously learning the optimal strategy for interacting with the environment.
[0057] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data are used to iteratively update the initial policy network to obtain a target policy network, including: constructing an electric energy constraint based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment and the electric load stored in the power storage equipment; wherein the first type of energy conversion equipment is a device that converts gas into electric energy; constructing a thermal energy constraint based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment and the heat load stored in the heat storage equipment; wherein the second type of energy conversion equipment is a device that converts gas into thermal energy; and iteratively updating the initial policy network based on the electric energy constraint, the thermal energy constraint, the constraint of the power storage equipment, the output power constraint of the energy supply system and the output power constraint of the energy storage system to obtain a target policy network.
[0058] In this first embodiment, the multiple constraints include at least energy balance constraints, power storage device constraints, energy supply system output power constraints, and energy storage system output power constraints. The energy balance constraints include electrical energy constraints and thermal energy constraints. The energy supply system output power constraints include output power constraints for the wind power system, photovoltaic system, and external power grid.
[0059] The power constraint can be expressed as formula 1:
[0060] P GT,t +P grid,t +P WT,t +P edis,t=P ech,t +L e (one)
[0061] Among them, P GT,t is the output power of the gas turbine (GT) at time t, P grid,t Represents the output power of the external power grid at time t, P WT,t represents the output power of the wind power system at time t, P edis,t is the discharge power of the storage device at time t, P ech,t is the charging power of the energy storage device at time t, L e The first type of energy conversion equipment is equipment that converts gas into electrical energy, such as a gas turbine.
[0062] The thermal energy constraint can be expressed as formula 2:
[0063]
[0064] Among them, P GB,t is the output power of the gas boiler (GB) at time t, P GT,t is the output power of the gas turbine (GT) at time t, a h Indicates thermal efficiency, that is, the actual heat output ratio per unit of electrical energy converted into heat energy, P hdis,t represents the discharge power of the heat storage device at time t, P hch,t is the charging power of the heat storage device at time t, L h The second type of energy conversion equipment is equipment that converts gas into thermal energy, such as gas boilers.
[0065] The output power constraints of the energy supply system can be shown as formulas 3 to 5.
[0066] P w_act ≤P w_max (three)
[0067] P pv_act ≤P pv_max (Four)
[0068]
[0069] P w_act Indicates the operating power of the wind power system, P w_max Indicates the maximum operating power of the wind system, P pv_act Indicates the operating power of the photovoltaic system, P pv_max Indicates the maximum operating power of the photovoltaic system, P grid,t represents the output power of the external power grid, Indicates the minimum output power of the external power grid, Indicates the maximum output power of the external power grid.
[0070] The output power constraint of the energy storage system can be expressed as formula 6:
[0071] -P po_max ≤P po (t)≤P po_max (six)
[0072] Among them, P po (t) represents the output power of the energy storage system at time t, P po_max Indicates the maximum output power of the energy storage system.
[0073] By establishing energy balance constraints, the park's energy supply precisely matches its needs, avoiding energy shortages or surpluses and maintaining stable operation of the park system. Establishing energy storage device constraints helps protect the energy storage system from overuse or extreme operating conditions, thereby extending its service life and reducing maintenance and replacement costs. By constraining the output power of the energy supply and storage systems, equipment operation is ensured to be within a safe and feasible range, avoiding the risks of overload or abnormal operation and improving the reliability and safety of the overall system.
[0074] By comprehensively considering multiple constraints, the scheduling strategy output by the intelligent agent is more in line with the actual operating environment, ensuring that the energy scheduling decisions proposed by the intelligent agent can be implemented in actual production and life, effectively guaranteeing the stable operation of the park's energy system, while improving the efficiency and economy of resource utilization, and achieving the goals of energy conservation, emission reduction and operating cost reduction.
[0075] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, the constraint conditions of the energy storage equipment are obtained by the following steps: constructing the charge state constraint conditions of the energy storage equipment based on the energy storage charge state of the energy storage equipment within a preset time period, the energy storage charging and discharging efficiency of the energy storage equipment, and the maximum energy storage capacity of the energy storage equipment; constructing the usage time constraint conditions of the energy storage equipment based on the charging power of the energy storage equipment, the discharging power of the energy storage equipment, the usage time limit of the energy storage equipment, and the maximum number of charge and discharge cycles of the energy storage equipment within the usage time limit within a preset time period; determining the constraint conditions of the energy storage equipment based on the usage time constraint conditions of the energy storage equipment and the charge state constraint of the energy storage equipment.
[0076] In the first embodiment, the state of charge constraint conditions of the power storage device can be shown as formulas 7 and 8:
[0077]
[0078] Among them, I SOC LFPB(t) represents the state of charge of the power storage system at time t (lithium iron phosphate battery can be abbreviated as LFPB, and the state of charge can be abbreviated as SOC), Indicates the minimum state of charge of the energy storage system, Indicates the maximum state of charge of the energy storage system, η LFPB Indicates the charging and discharging efficiency of the energy storage system, E LFPB,max Indicates the maximum capacity of the power storage system, Indicates the initial state of charge of the energy storage system.
[0079] The usage time constraint of the energy storage device can be shown as formula 9:
[0080]
[0081] Wherein, Pd(t) represents the discharge power of the energy storage system at time t, Pc(t) represents the charging power of the energy storage system at time t, Cmax represents the maximum number of charge and discharge cycles within the service life of the energy storage system, and YEOL represents the service life of the energy storage system (which can be measured in years).
[0082] By establishing usage time constraints, the number of charge and discharge cycles of the energy storage device within a specific time period is limited, avoiding battery aging caused by frequent charging and discharging, thereby significantly extending the service life of the energy storage device and reducing long-term operating costs. By establishing state of charge constraints, the SOC of the energy storage device is ensured to remain within a safe range, preventing equipment shutdown caused by too low SOC or overcharging risks caused by too high SOC, ensuring the stable operation of the energy storage system under high safety standards.
[0083] Through the above steps, the energy storage equipment is protected at the physical level, avoiding potential safety hazards and economic losses. At the same time, it enables the intelligent body to more intelligently schedule the charging and discharging of the energy storage equipment, taking advantage of external energy price fluctuations and the load demand of the park to select the best time for charging and discharging, thereby improving the efficiency of energy scheduling.
[0084] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, a target value network is constructed based on the reward function, and an initial strategy network is constructed based on the target value network, including: determining a first loss function based on the reward function, state variables, action variables corresponding to the state variables, and expected return information; calculating the network parameters of the target value network during iterative updates by performing gradient operations on the first loss function, and constructing the target value network; determining a second loss function of the initial strategy network based on the maximum value of the target value network; calculating the network parameters of the initial strategy network during iterative updates by performing gradient operations on the second loss function, and constructing the initial strategy network.
[0085] In the first embodiment, the first loss function can be shown as formulas 10 and 11,
[0086] L(θ Q )=E[y t -Q(s t ,a t |θ Q )] 2 (eleven)
[0087] y t =r t +γQ′[s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ] (twelve)
[0088] Among them, s t Represents the state variable, a t Represents the state variable s t The corresponding action variable, y t The target Q value (i.e., the expected return information mentioned above), that is, the value of the Q function, is used to measure the expected cumulative reward that can be obtained under a specific state variable st-action variable at pair in the reinforcement learning algorithm. Q(s t ,a t |θ Q ) is the Q value output by the current value network at time ι, r t is the instantaneous reward at time t extracted from multiple sample data (i.e., the instantaneous reward in the reward function); π′(s t+1 |θ π′ ) is the target policy network with parameters θ π′ Next input state variable s t+1 The action variable output when ; For the target value network in parameters Next input state s t+1 and action variables π′(s t+1 |θ τ′ ) under the input Q value. According to the gradient update rule, by (i.e. the first loss function mentioned above) to find the gradient and obtain the update formula of the value network (i.e. the target value network mentioned above), which can be shown as the following formula 13:
[0089]
[0090] in, is the value network parameter during the k-th round of learning; μ Q is the learning rate of the value network; is the loss function L(θ k -1 Q ) for parameters gradient.
[0091] The policy network needs to learn to maximize the Q value output by the value network. Therefore, the output Q function of the value network can be used as the loss function of the policy network (i.e., the second loss function mentioned above). By calculating the policy gradient of the Q function, the update formula of the policy network can be obtained as shown in Formula 14.
[0092]
[0093] in, is the current policy network parameter at the kth round of learning; μ π is the learning rate of the policy network; is the policy gradient. In order to ensure the stability of the learning process, a soft update technique is usually adopted for the target policy network. The calculation formula for slowly updating the parameters of the target policy network (i.e. the initial policy network mentioned above) can be shown as Formula 15.
[0094]
[0095] in, and are the target value network and target strategy network parameters in the kth round of learning respectively; τ is the soft update coefficient.
[0096] By learning the reward function, the target value network (Q network) can predict the expected cumulative reward under different state variable-action variable pairs, providing a basis for the agent's decision-making and improving the reliability of the agent's decision-making. By constructing an initial policy network based on the reward function and interactively learning with the value network, it can quickly adjust and optimize its own strategy, effectively shortening the time from blind trial and error to strategy maturity. Through techniques such as soft updates, the stability of the learning process is guaranteed, avoiding common training issues such as difficulty in policy convergence caused by unstable reward signals.
[0097] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, an objective function is constructed based on the operating costs of multiple energy systems and the total carbon emissions of the park, including: determining the gas purchase cost of the park based on the gas purchase price of the park, the output power of the first type of energy conversion equipment, and the conversion efficiency of the third type of energy conversion equipment; wherein the third type of energy conversion equipment represents equipment that converts gas into electrical energy or thermal energy; determining the energy storage equipment in the energy storage system, and determining the operating cost of the energy storage system based on the number of energy storage equipment and the unit operating cost of the energy storage equipment; determining the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; and constructing an objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid, and the carbon emission coefficient of the first type of energy conversion equipment.
[0098] In this embodiment 1, the objective function can be determined based on the park operation cost and the park environmental cost (ie, carbon emissions). The objective function can be shown in Formula 16:
[0099] min C G =μ1C run +μ2C carbon (sixteen)
[0100] Among them, CG represents the comprehensive cost of the park, Crun represents the system operation cost of the park, Ccarbon represents the environmental cost of the park, that is, carbon emissions, and u1 and u2 are the weight factors of the two optimization indicators respectively.
[0101] The park operating cost can be determined based on the operating cost of the energy storage system, the park's electricity purchase cost, and the park's gas purchase cost. The calculation formula for the park operating cost can be shown in Formula 17:
[0102]
[0103] Among them, Cstorage represents the operating cost of the energy storage system, Ceb represents the electricity purchase cost of the park, Cgb represents the gas purchase cost of the park, and ε grid and ε gas are electricity purchase price, gas purchase price, P GT,t and P GB,t are the output power of gas turbine (GT) and gas boiler (GB) at time t, η GT and η GB are the conversion efficiencies of gas turbine (GT) and gas boiler (GB), P grid The third category of energy conversion equipment refers to equipment that converts gas into electricity or heat, such as gas turbines and gas boilers.
[0104] The operating cost of the energy storage system can be determined based on the output power, working hours and unit operating cost of the energy storage system. The calculation formula for the operating cost of the energy storage system can be shown in Formula 18:
[0105] R cos =ΣV E |P po (t)|Δt (XVIII)
[0106] Among them, VE represents the unit operating cost of the energy storage system, P po (t) represents the output power of the energy storage system at time t, and Δt represents the unit time.
[0107] The calculation formula of the park's environmental cost can be shown in Formula 19:
[0108]
[0109] Among them, Fd t CE , Fd t GT and Fd t GB are the carbon emission coefficients for electricity purchase, gas turbine (GT) and gas boiler (GB) operation, respectively.
[0110] By integrating electricity and gas purchase costs with energy storage operating expenses, the objective function guides the agent to find the lowest-cost energy combination while ensuring energy supply, significantly reducing the park's energy operating expenses. Furthermore, by introducing carbon emission costs, the agent's energy scheduling decisions prioritize low-carbon or carbon-free energy sources such as photovoltaic and wind power, efficiently utilize energy storage systems, and reduce reliance on fossil fuels, thereby reducing the park's carbon emissions.
[0111] Optionally, in the energy scheduling method within the park provided in Example 1 of the present application, the reward function includes at least: an instant reward function and a cumulative reward function; the reward function is constructed based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data, including: constructing state variables based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge status information of the storage device, the output power prediction value of the wind system, and the output power prediction value of the photovoltaic system; constructing action variables based on the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; constructing a reward function for energy scheduling of the park at any time based on the state variables, action variables and objective function to obtain an instant reward function; constructing a cumulative reward function based on the instant reward function and preset parameters of the energy scheduling of the park within the preset time period.
[0112] In this first embodiment, during the reinforcement learning process, the agent continuously iterates its strategy through interaction with the environment, gradually generating a better strategy to achieve the optimization goal. To transform the existing conventional optimization model into a reinforcement learning model, it is necessary to define the state space, action space, and reward function of the campus energy system.
[0113] Specifically, the electric load, thermal load, charging and discharging power of energy storage, SOC, wind power and photovoltaic power forecast output during the park scheduling period can be selected as the state space, as shown in the following formula 20:
[0114] S={S EL ,S TL ,S bt ,S soc ,S wt ,S pv}(twenty)
[0115] Among them, SEL is the electric load during the park scheduling period; STL is the electric load, Sbt and Ssoc are the charging and discharging power and state of charge (SOC) of the energy storage system respectively; Swt and Spv are the predicted output power of the wind power system and photovoltaic system respectively.
[0116] Then, the model's decision variables are usually selected as the action space of the park energy system, such as the output power of wind power, photovoltaic power, and energy storage. However, directly selecting output power makes it difficult to express coupling. To simplify the complexity of model learning and consider the temporal coupling between decision variables, the output power increments of energy storage, gas turbines, and gas boiler equipment are selected as the action space, as shown in the following formula 21:
[0117] A={A bt ,A GT ,A CB ,A P2G ,A EB}(Twenty-one)
[0118] Among them, Abt, AGT, AGB, AP2G, and AEB are the output power increments of the energy storage system, gas turbine, gas boiler, power-to-gas equipment, and electric boiler, respectively.
[0119] Secondly, in order to train the agent to learn the scheduling strategy that minimizes the total cost of the park energy system, the negative value of the objective function can be set as the reward function. Therefore, based on reinforcement learning theory, it is necessary to transform the mathematical problem of minimization into a form of maximizing the reward value to guide the agent's strategy learning. That is, the smaller the total cost, the greater the reward. The calculation formula of the immediate reward rt can be shown in Formula 22:
[0120] r t =-K(μ1C run +μ2Ccarbon )(Twenty-two)
[0121] Where rt is the immediate reward obtained by the agent when it selects the action variable at = [Awt, Apv, Abt, AGT, AGB] under any state variable st = [SEL, STL, Sbt, Ssoc, Swt, Spv}, and K is the adjustment coefficient. The cumulative reward function R can be shown as formula 23:
[0122]
[0123] Among them, R is the cumulative reward obtained by the agent based on the scheduling plan corresponding to the state variable; γ is the discount factor, which indicates the importance of the rewards in the future time period relative to the current moment. For example, when γ = 0, it means that only the current immediate reward is considered without considering the future long-term reward. When γ = 1, it means that the future long-term reward and the current immediate reward are equally important.
[0124] By integrating real-time park energy load, energy storage charging and discharging power, state of charge (SOC), and predicted wind and photovoltaic power output, the intelligent agent can comprehensively and accurately perceive the current energy environment, providing a solid data foundation for efficient decision-making. By designing action variables, the intelligent agent can directly control the changing trends of different energy systems, achieving dynamic response and instantly adjusting energy allocation strategies to cope with load fluctuations and the uncertainty of renewable energy.
[0125] By constructing a cumulative reward function that incorporates future rewards into current decisions as a discount factor, the intelligent agent is encouraged to prioritize energy scheduling beyond immediate cost savings or carbon emissions reductions, taking into account long-term economic and environmental benefits. This ensures forward-looking and holistic decision-making. The immediate reward function ensures that the agent receives feedback at every decision moment. It rewards the agent based on the immediate impact of current actions on economic costs and carbon emissions, while the cumulative reward function focuses on long-term cumulative effects. The combination of the two helps the agent find the optimal balance between short-term behavior and long-term goals.
[0126] Optionally, in this embodiment 1, Figure 2This is a flowchart of a method for optimizing industrial park economic carbon reduction based on a deep reinforcement learning algorithm. First, considering the complex interactions among various energy loads, such as electricity, heat, and gas, as well as the supply, conversion, and storage sides of the industrial park energy system, a mathematical model is designed and established to simultaneously minimize the industrial park's economic costs and carbon emissions, namely the objective function mentioned above. This objective function is then transformed into a deep reinforcement learning framework, which includes defining the interaction between the agent and the environment (the industrial park energy system), clarifying key elements such as the state space (such as energy load information and energy storage status), the action space (such as device output power adjustment), and the reward function (based on economic costs and carbon emissions). The model is then solved using the Deep Deterministic Policy Gradient (DDPG) algorithm. The DDPG algorithm is capable of handling complex optimization problems in continuous action spaces. Through coordinated learning of the value network and policy network, it gradually optimizes the agent's policy to achieve optimal industrial park energy scheduling. Finally, through iterative learning and optimization, the agent's policy is continuously adjusted to adapt to the dynamic changes of the industrial park energy system. This process involves training the agent to output the optimal equipment scheduling plan based on real-time status data, minimizing both economic costs and carbon emissions. The optimization results will guide the actual operation of the park's energy system, dynamically adjusting the output of energy equipment to achieve the optimal balance between economic and environmental protection.
[0127] Optionally, in this embodiment 1, Figure 3 This is a schematic diagram of the park's comprehensive energy optimization and dispatching system. The energy supply side integrates wind power, photovoltaic power, and the external grid to provide multi-type energy input. The energy conversion side uses equipment such as P2G (power-to-gas), GT (gas turbines), GB (gas boilers), and EB (electric boilers) to achieve the conversion and coupling of electricity, heat, and gas. The energy load side covers gas load, heat load, and electricity load, representing the actual demand for various energy sources within the park.
[0128] Figure 3 The diagram also shows the energy conversion relationships and flow directions. Gas can be supplied directly to gas loads, converted to electricity through GTs, or generated through GBs to generate heat. Gas storage tanks act as a buffer, ensuring flexible gas allocation. Electricity can be supplied directly to electrical loads through GTs and the external grid, or generated through EBs to be used by GBs. Storage facilities help regulate peak and valley power levels. Heat can be supplied directly to thermal loads by GTs, GBs, and EBs.
[0129] Optionally, in this embodiment 1, Figure 4This is a diagram of the campus structure based on the Markov Decision Process (MDP). Agent is the executor of this solution, that is, it interacts with the environment and decides what actions to take based on the received state information. Environment represents the scenario where the agent is located and needs to make decisions, such as the above-mentioned campus integrated energy system, which includes all factors that may affect the decision. State (state, that is, the state variable mentioned above) represents the observable state information in the environment. Figure 4 S t (Current Status) and S t+1 (Next state). The state covers key data such as the energy load, energy storage status, energy forecast, etc. of the park. The agent relies on this information to make decisions. Action (action, that is, the action variable mentioned above) is the behavior taken by the agent in a given state, such as adjusting the charging and discharging power of the power storage facility or the output of the gas turbine. The action directly affects the change of the environmental state. Reward is the feedback received from the environment after the agent performs the action, which is used to evaluate the effect of its action and guide the agent to pursue the best energy scheduling strategy. Figure 4 Marked as r t (immediate reward) and r t+1 (Subsequent rewards). The agent is based on the current state S t Take action, and the environment enters a new state S t+1 , and give corresponding rewards t This cycle of state transition and reward feedback prompts the agent to continuously improve its strategy through learning.
[0130] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0131] Example 2
[0132] The second embodiment of the present application further provides an energy scheduling device within a park. It should be noted that the energy scheduling device within the park of the second embodiment of the present application can be used to execute the energy scheduling method for the park provided in the first embodiment of the present application. The energy scheduling device within the park provided in the second embodiment of the present application is introduced below.
[0133] Figure 5 This is a schematic diagram of an energy dispatching device in a park according to the second embodiment of the present application. Figure 5 As shown, the device includes: a collection unit 501, a calculation unit 502 and a scheduling unit 503.
[0134] Specifically, the collection unit 501 is used to collect the operating status of multiple energy systems in the park during the target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system. The energy supply system includes at least: an external power grid, a wind power system, and a photovoltaic system. The energy storage system includes at least: electricity storage equipment, gas storage equipment, and heat storage equipment.
[0135] The computing unit 502 is used to input the state variables into the target policy network based on the deep deterministic policy gradient algorithm and output the action variables, wherein the action variables are used to represent the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and the target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network.
[0136] The scheduling unit 503 is used to perform energy scheduling on multiple energy systems in the park according to the action variables.
[0137] The energy scheduling device in the park provided in the second embodiment of the present application uses a collection unit 501 to collect the operating status of multiple energy systems in the park during a target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system. The energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system. The energy storage system includes at least: an electric storage device, a gas storage device, and a heat storage device. The calculation unit 502 inputs the state variables into the target policy network based on the deep deterministic policy gradient algorithm and outputs action variables, wherein the action variables are used to represent the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the park energy scheduling under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network, and the scheduling unit 503 performs energy scheduling for the multiple energy systems in the park based on the action variables. This solves the problem in the related art that due to the existence of multiple subjects in the park, the objective function corresponding to the energy scheduling of each subject and the constraints interact with each other, making it difficult to find a global optimal solution for the optimization problem, resulting in poor energy conservation and carbon reduction effects in the park.
[0138] By monitoring and recording the power, wind, and photovoltaic supply within the park, as well as the dynamics of electricity, gas, and heat storage, a data set reflecting the overall status of the energy system is established, enabling the intelligent agent to make more accurate decisions based on current and predicted energy demand and supply. The target policy network is trained through the DDPG algorithm, which can process continuous action space and effectively respond to the dynamic balance and complex coupling relationships within the system. Even in the face of unexpected load fluctuations or changes in energy prices, the system can maintain efficient operation. By setting multiple constraints, it is ensured that the action variables output by the intelligent agent can not only reduce costs and emissions, but also comply with physical and operational limitations, avoiding unrealistic or harmful scheduling decisions. Through interaction with the target value network, the intelligent agent continuously learns and optimizes its strategy to obtain the highest return (that is, the above-mentioned return information), that is, the minimum total cost and carbon emissions.
[0139] In summary, the agent's decision-making has achieved dual optimization of the park's energy system's economic efficiency and carbon reduction goals. It not only responds to changes in the park's energy demand in real time, but also predicts and adapts to market and weather uncertainties, reducing both operating costs and carbon emissions. The solution demonstrates its flexibility and robustness across diverse scenarios, providing the park with a stable, economical, and environmentally friendly energy supply and significantly enhancing the intelligent level of energy management.
[0140] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned device also includes: a first construction unit, used to construct an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park; a second construction unit, used to construct a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; a third construction unit, used to construct a target value network based on the reward function, and to construct an initial policy network based on the target value network; an update unit, used to iteratively update the initial policy network using a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data to obtain a target policy network.
[0141] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned update unit includes: a first construction subunit, used to construct an electric energy constraint condition based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment, and the electric load stored in the power storage equipment; wherein, the first type of energy conversion equipment is a device that converts gas into electric energy; a second construction subunit, used to construct a thermal energy constraint condition based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment, and the heat load stored in the heat storage equipment; wherein, the second type of energy conversion equipment is a device that converts gas into thermal energy; an update subunit, used to iteratively update the initial strategy network based on the electric energy constraint condition, the thermal energy constraint condition, the constraint condition of the power storage equipment, the output power constraint condition of the energy supply system, and the output power constraint condition of the energy storage system to obtain a target strategy network.
[0142] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned update unit also includes: a third construction sub-unit, used to construct a state of charge constraint condition for the energy storage device based on the energy storage state of charge of the energy storage device within a preset time period, the energy storage charging and discharging efficiency of the energy storage device, and the maximum energy storage capacity of the energy storage device; a fourth construction sub-unit, used to construct a usage time constraint condition for the energy storage device based on the charging power of the energy storage device, the discharging power of the energy storage device, the usage time limit of the energy storage device, and the maximum number of charge and discharge cycles of the energy storage device within the usage time limit within a preset time period; a first determination sub-unit, used to determine the constraint condition of the energy storage device based on the usage time constraint condition of the energy storage device and the state of charge constraint of the energy storage device.
[0143] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned third construction unit includes: a second determination subunit, used to determine the first loss function based on the reward function, state variables, action variables corresponding to the state variables, and expected return information; a first calculation subunit, used to calculate the network parameters of the target value network during iterative update by performing gradient operation on the first loss function, and construct the target value network; the third determination subunit, used to determine the second loss function of the initial policy network based on the maximum value of the target value network; the second calculation subunit, used to calculate the network parameters of the initial policy network during iterative update by performing gradient operation on the second loss function, and construct the initial policy network.
[0144] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned first construction unit includes: a fourth determination subunit, used to determine the gas purchase cost of the park based on the gas purchase price of the park and the output power of the third type of energy conversion equipment and the conversion efficiency of the first type of energy conversion equipment; wherein the third type of energy conversion equipment refers to equipment that converts gas into electrical energy or thermal energy; a fifth determination subunit, used to determine the energy storage equipment in the energy storage system, and determine the operating cost of the energy storage system based on the number of energy storage equipment and the unit operating cost of the energy storage equipment; a sixth determination subunit, used to determine the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; the fifth construction subunit, used to construct an objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park and the gas purchase cost of the park, the carbon emission coefficient of the external power grid and the carbon emission coefficient of the first type of energy conversion equipment.
[0145] Optionally, in the energy scheduling device within the park provided in Example 2 of the present application, the above-mentioned reward function includes at least: an instant reward function and a cumulative reward function; the second construction unit includes: a sixth construction sub-unit, which is used to construct state variables based on the energy load information of the park within a preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge status information of the storage device, the output power prediction value of the wind system, and the output power prediction value of the photovoltaic system; a seventh construction sub-unit, which is used to construct action variables based on the output power increment of the energy storage system and the output power increment of the energy conversion system within a preset time period; an eighth construction sub-unit, which is used to construct a reward function for energy scheduling of the park at any time based on the state variables, action variables and objective function to obtain an instant reward function; a ninth construction sub-unit, which is used to construct a cumulative reward function based on the instant reward function and preset parameters of the energy scheduling of the park within a preset time period.
[0146] The energy scheduling device in the park includes a processor and a memory. The above-mentioned acquisition unit 501, calculation unit 502 and scheduling unit 503 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0147] The processor contains a core, which retrieves the corresponding program unit from the memory. One or more cores can be set, and the energy conservation and carbon reduction effects within the park can be improved by adjusting the core parameters.
[0148] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0149] A third embodiment of the present invention provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a method for energy scheduling within a park is implemented.
[0150] A fourth embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes an energy scheduling method within a park when running.
[0151] like Figure 6 As shown, embodiment 5 of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and can be run on the processor. When the processor executes the program, the following steps are implemented: collecting the operating status of multiple energy systems in the park within a target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system, and the energy storage system includes at least: power storage equipment, gas storage equipment, and heat storage equipment; based on a deep deterministic policy gradient algorithm, the state variables are input into a target policy network, and action variables are output, wherein the action variables are used to characterize the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network; and energy scheduling is performed on multiple energy systems in the park based on the action variables.
[0152] When the processor executes the program, it also implements the following steps: the target policy network is obtained by iterative training of the following steps: constructing an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park; constructing a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; constructing a target value network based on the reward function, and constructing an initial policy network based on the target value network; using a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data to iteratively update the initial policy network to obtain a target policy network.
[0153] When the processor executes the program, it also implements the following steps: using a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data to iteratively update the initial policy network to obtain a target policy network, including: constructing an electric energy constraint based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment and the electric load stored by the power storage equipment; wherein the first type of energy conversion equipment is a device that converts gas into electric energy; constructing a thermal energy constraint based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment and the heat load stored by the heat storage equipment; wherein the second type of energy conversion equipment is a device that converts gas into thermal energy; iteratively updating the initial policy network based on the electric energy constraint, the thermal energy constraint, the constraint of the power storage equipment, the output power constraint of the energy supply system and the output power constraint of the energy storage system to obtain a target policy network.
[0154] When the processor executes the program, the following steps are also implemented: the constraints of the power storage device are obtained by the following steps: the state of charge constraints of the power storage device are constructed based on the energy storage state of charge of the power storage device within a preset time period, the energy storage charging and discharging efficiency of the power storage device, and the maximum energy storage capacity of the power storage device; the usage time constraints of the power storage device are constructed based on the charging power of the power storage device within a preset time period, the discharging power of the power storage device, the usage time limit of the power storage device, and the maximum number of charge and discharge cycles of the power storage device within the usage time limit; the constraints of the power storage device are determined based on the usage time constraints of the power storage device and the state of charge constraints of the power storage device.
[0155] When the processor executes the program, it also implements the following steps: constructing a target value network based on the reward function, and constructing an initial policy network based on the target value network, including: determining a first loss function based on the reward function, state variables, action variables corresponding to the state variables, and expected return information; calculating the network parameters of the target value network during iterative updates by performing gradient operations on the first loss function, and constructing the target value network; determining a second loss function of the initial policy network based on the maximum value of the target value network; calculating the network parameters of the initial policy network during iterative updates by performing gradient operations on the second loss function, and constructing the initial policy network.
[0156] When the processor executes the program, it also implements the following steps: constructing an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park, including: determining the gas purchase cost of the park based on the gas purchase price of the park and the output power of the first type of energy conversion equipment and the conversion efficiency of the third type of energy conversion equipment; wherein the third type of energy conversion equipment represents equipment that converts gas into electrical energy or thermal energy; determining the energy storage equipment in the energy storage system, and determining the operating cost of the energy storage system based on the number of energy storage equipment and the unit operating cost of the energy storage equipment; determining the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; constructing an objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid, and the carbon emission coefficient of the first type of energy conversion equipment.
[0157] When the processor executes the program, it also implements the following steps: the reward function includes at least: an instant reward function and a cumulative reward function; the reward function is constructed based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data, including: constructing state variables based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge state information of the storage device, the output power prediction value of the wind system, and the output power prediction value of the photovoltaic system; constructing action variables based on the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; constructing a reward function for the energy scheduling of the park at any time based on the state variables, action variables and objective function to obtain an instant reward function; constructing a cumulative reward function based on the instant reward function and preset parameters for the energy scheduling of the park within the preset time period.
[0158] The devices in this article can be servers, PCs, PADs, mobile phones, etc.
[0159] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: collecting the operating status of multiple energy systems in a park within a target time period to obtain state variables, wherein the multiple energy systems include at least: an energy supply system, an energy conversion system, and an energy storage system, the energy supply system includes at least: an external power grid, a wind system, and a photovoltaic system, and the energy storage system includes at least: power storage equipment, gas storage equipment, and heat storage equipment; based on a deep deterministic policy gradient algorithm, the state variables are input into a target policy network and action variables are output, wherein the action variables are used to characterize the action decision information of the park scheduling multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and a target value network; the initial policy network is responsible for outputting the action decision information of the park's energy scheduling under any state variables; the target value network is used to predict the return information of the action variables output by the policy network; and energy scheduling is performed on multiple energy systems in the park based on the action variables.
[0160] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the target policy network is obtained by iterative training by the following steps: constructing an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park; constructing a reward function based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data; constructing a target value network based on the reward function, and constructing an initial policy network based on the target value network; using a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data to iteratively update the initial policy network to obtain a target policy network.
[0161] When executed on a data processing device, it is also suitable for executing a program that is initialized with the following method steps: using a deep deterministic policy gradient algorithm, multiple constraints and multiple sample data to iteratively update the initial policy network to obtain a target policy network, including: constructing an electric energy constraint based on the output power of the wind system, the output power of the first type of energy conversion equipment, the output power of the external power grid, the discharge power of the power storage equipment and the electric load stored by the power storage equipment; wherein the first type of energy conversion equipment is a device that converts gas into electric energy; constructing a thermal energy constraint based on the output power of the second type of energy conversion equipment, the charging power of the heat storage equipment, the discharge power of the heat storage equipment and the heat load stored by the heat storage equipment; wherein the second type of energy conversion equipment is a device that converts gas into thermal energy; iteratively updating the initial policy network based on the electric energy constraint, the thermal energy constraint, the constraint of the power storage equipment, the output power constraint of the energy supply system and the output power constraint of the energy storage system to obtain a target policy network.
[0162] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the constraints of the power storage device are obtained by the following steps: constructing the state of charge constraints of the power storage device based on the energy storage state of charge of the power storage device within a preset time period, the energy storage charging and discharging efficiency of the power storage device, and the maximum energy storage capacity of the power storage device; constructing the usage time constraints of the power storage device based on the charging power of the power storage device, the discharging power of the power storage device, the usage time limit of the power storage device, and the maximum number of charge and discharge cycles of the power storage device within the usage time limit within a preset time period; and determining the constraints of the power storage device based on the usage time constraints of the power storage device and the state of charge constraints of the power storage device.
[0163] When executed on a data processing device, it is also suitable for executing an initialization program with the following method steps: constructing a target value network based on a reward function, and constructing an initial policy network based on the target value network, including: determining a first loss function based on the reward function, state variables, action variables corresponding to the state variables, and expected return information; calculating the network parameters of the target value network during iterative updates by performing gradient operations on the first loss function, and constructing the target value network; determining the second loss function of the initial policy network based on the maximum value of the target value network; calculating the network parameters of the initial policy network during iterative updates by performing gradient operations on the second loss function, and constructing the initial policy network.
[0164] When executed on a data processing device, it is also suitable for executing an initialized program having the following method steps: constructing an objective function based on the operating costs of multiple energy systems and the total carbon emissions of the park, including: determining the gas purchase cost of the park based on the gas purchase price of the park and the output power of the first type of energy conversion equipment and the conversion efficiency of the third type of energy conversion equipment; wherein the third type of energy conversion equipment represents equipment that converts gas into electrical energy or thermal energy; determining the energy storage equipment in the energy storage system, and determining the operating cost of the energy storage system based on the number of energy storage equipment and the unit operating cost of the energy storage equipment; determining the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; constructing an objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park and the gas purchase cost of the park, the carbon emission coefficient of the external power grid and the carbon emission coefficient of the first type of energy conversion equipment.
[0165] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the reward function includes at least: an instant reward function and a cumulative reward function; the reward function is constructed based on the objective function, state variables corresponding to multiple sample data within a preset time period, and action variables corresponding to multiple sample data, including: constructing state variables based on the energy load information of the park within a preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the charge state information of the storage device, the output power prediction value of the wind system, and the output power prediction value of the photovoltaic system; constructing action variables based on the output power increment of the energy storage system and the output power increment of the energy conversion system within a preset time period; constructing a reward function for energy scheduling of the park at any time based on the state variables, action variables and objective function to obtain an instant reward function; constructing a cumulative reward function based on the instant reward function and preset parameters of energy scheduling of the park within a preset time period.
[0166] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0168] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0170] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0171] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0172] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0173] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0174] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0175] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for energy scheduling within a park, characterized in that: include: Collecting the operating status of multiple energy systems in the park during a target time period to obtain state variables, wherein the multiple energy systems include at least an energy supply system, an energy conversion system, and an energy storage system. The energy supply system includes at least an external power grid, a wind power system, and a photovoltaic system. The energy storage system includes at least an electricity storage device, a gas storage device, and a heat storage device. Based on the deep deterministic policy gradient algorithm, state variables are input into the target policy network, and action variables are output, wherein the action variables are used to represent the action decision information of the park scheduling the multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network according to multiple constraints and the target value network; the initial policy network is responsible for outputting the action decision information of the energy scheduling of the park under any state variables; the target value network is used to predict the reward information of the action variables output by the policy network; Energy scheduling is performed on multiple energy systems in the park according to the action variables.
2. The method according to claim 1, characterized in that The target policy network is iteratively trained by the following steps: Constructing an objective function based on the operating costs of the multiple energy systems and the total carbon emissions of the park; Constructing a reward function based on the objective function, state variables corresponding to a plurality of sample data within a preset time period, and action variables corresponding to a plurality of sample data; Constructing the target value network according to the reward function, and constructing an initial policy network according to the target value network; The deep deterministic policy gradient algorithm, the multiple constraints, and the multiple sample data are used to iteratively update the initial policy network to obtain the target policy network.
3. The method according to claim 2, characterized in that Iteratively updating the initial policy network using the deep deterministic policy gradient algorithm, the multiple constraints, and the multiple sample data to obtain the target policy network includes: An electric energy constraint condition is established based on the output power of the wind power system, the output power of the first type of energy conversion device, the output power of the external power grid, the discharge power of the power storage device, and the electric load stored by the power storage device; wherein the first type of energy conversion device is a device that converts gas into electric energy; Establishing a thermal energy constraint condition based on the output power of the second type of energy conversion device, the charging power of the heat storage device, the discharging power of the heat storage device, and the heat load stored by the heat storage device; wherein the second type of energy conversion device is a device that converts gas into thermal energy; The initial strategy network is iteratively updated according to the electric energy constraint condition, the thermal energy constraint condition, the constraint condition of the power storage device, the output power constraint condition of the energy supply system and the output power constraint condition of the energy storage system to obtain the target strategy network.
4. The method according to claim 3, characterized in that The constraints of the power storage device are obtained by the following steps: Establishing a state of charge constraint condition for the power storage device according to the energy storage state of charge of the power storage device within a preset time period, the energy storage charging and discharging efficiency of the power storage device, and the maximum energy storage capacity of the power storage device; Establishing a usage time constraint condition for the power storage device based on the charging power of the power storage device, the discharging power of the power storage device, the usage time limit of the power storage device, and the maximum number of charge and discharge cycles of the power storage device within the usage time limit within a preset time period; The constraint condition of the electric storage device is determined according to the usage time constraint condition of the electric storage device and the charge state constraint of the electric storage device.
5. The method according to claim 2, characterized in that Constructing the target value network according to the reward function, and constructing the initial policy network according to the target value network, including: Determining a first loss function based on the reward function, the state variable, the action variable corresponding to the state variable, and expected return information; By performing a gradient operation on the first loss function, network parameters of the target value network during iterative update are calculated to construct the target value network; Determining a second loss function of the initial strategy network according to the maximum value of the target value network; By performing a gradient operation on the second loss function, network parameters of the initial policy network during iterative update are calculated to construct the initial policy network.
6. The method according to claim 2, characterized in that An objective function is constructed based on the operating costs of the multiple energy systems and the total carbon emissions of the park, including: The gas purchase cost of the park is determined based on the gas purchase price of the park, the output power of the first type of energy conversion equipment, and the conversion efficiency of the third type of energy conversion equipment; wherein the third type of energy conversion equipment refers to equipment that converts gas into electrical energy or thermal energy; Determining energy storage devices in the energy storage system, and determining an operating cost of the energy storage system based on the number of the energy storage devices and the unit operating costs of the energy storage devices; Determining the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; The objective function is constructed based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid, and the carbon emission coefficient of the first type of energy conversion equipment.
7. The method according to claim 2, characterized in that The reward function includes at least: an instantaneous reward function and a cumulative reward function; the reward function is constructed based on the objective function, state variables corresponding to a plurality of sample data within a preset time period, and action variables corresponding to a plurality of sample data, including: Constructing the state variable based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the state of charge information of the power storage device, the output power prediction value of the wind power system, and the output power prediction value of the photovoltaic system; Constructing the action variable according to the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; Constructing a reward function for the energy scheduling of the park at any time based on the state variable, the action variable, and the objective function to obtain the instantaneous reward function; The cumulative reward function is constructed based on the instant reward function and preset parameters of the park energy scheduling within a preset time period.
8. An energy dispatching device in a park, characterized in that: include: a collection unit, configured to collect operating status information of multiple energy systems within the park during a target time period to obtain state variables, wherein the multiple energy systems include at least an energy supply system, an energy conversion system, and an energy storage system; the energy supply system includes at least an external power grid, a wind power system, and a photovoltaic system; and the energy storage system includes at least an electricity storage device, a gas storage device, and a heat storage device; A computing unit is configured to input state variables into a target policy network based on a deep deterministic policy gradient algorithm and output action variables, wherein the action variables are used to represent action decision information for the park scheduling the multiple energy systems; the target policy network is a network model obtained by iteratively gradient updating the initial policy network based on multiple constraints and a target value network; the initial policy network is responsible for outputting action decision information for the energy scheduling of the park under any state variables; and the target value network is used to predict reward information for the action variables output by the policy network; A scheduling unit is used to perform energy scheduling on multiple energy systems in the park according to the action variables.
9. The device according to claim 8, characterized in that The device further comprises: A first construction unit is configured to construct an objective function based on the operating costs of the multiple energy systems and the total carbon emissions of the park; A second construction unit is configured to construct a reward function based on the objective function, state variables corresponding to a plurality of sample data within a preset time period, and action variables corresponding to a plurality of sample data; A third construction unit is configured to construct the target value network according to the reward function, and to construct an initial policy network according to the target value network; An updating unit is configured to iteratively update the initial policy network using the deep deterministic policy gradient algorithm, the multiple constraints, and the multiple sample data to obtain the target policy network.
10. The device according to claim 9, characterized in that The updating unit includes: a first constructing subunit, configured to construct an electric energy constraint condition based on the output power of the wind power system, the output power of the first type of energy conversion device, the output power of the external power grid, the discharge power of the power storage device, and the electric load stored by the power storage device; wherein the first type of energy conversion device is a device that converts gas into electric energy; a second construction subunit, configured to construct a thermal energy constraint condition based on the output power of a second type of energy conversion device, the charging power of the heat storage device, the discharging power of the heat storage device, and the heat load stored by the heat storage device; wherein the second type of energy conversion device is a device that converts gas into thermal energy; An updating subunit is used to iteratively update the initial strategy network based on the electric energy constraint condition, the thermal energy constraint condition, the constraint condition of the power storage device, the output power constraint condition of the energy supply system, and the output power constraint condition of the energy storage system to obtain the target strategy network.
11. The device according to claim 10, characterized in that The updating unit includes: A third construction subunit is used to construct a state of charge constraint condition of the power storage device according to the energy storage state of charge of the power storage device within a preset time period, the energy storage charging and discharging efficiency of the power storage device, and the maximum energy storage capacity of the power storage device; A fourth construction subunit is configured to construct a usage time constraint condition for the power storage device based on the charging power of the power storage device, the discharging power of the power storage device, the usage time limit of the power storage device, and the maximum number of charge and discharge cycles of the power storage device within the usage time limit within a preset time period; The first determining subunit is configured to determine a constraint condition of the electric storage device according to a usage time constraint condition of the electric storage device and a charge state constraint condition of the electric storage device.
12. The device according to claim 9, characterized in that The third building block comprises: A second determining subunit is configured to determine a first loss function based on the reward function, the state variable, the action variable corresponding to the state variable, and the expected return information; A first calculation subunit is configured to calculate network parameters of the target value network during iterative update by performing a gradient operation on the first loss function, thereby constructing the target value network; A third determining subunit is configured to determine a second loss function of the initial strategy network according to the maximum value of the target value network; The second computing subunit is configured to calculate network parameters of the initial policy network during iterative update by performing a gradient operation on the second loss function, thereby constructing the initial policy network.
13. The device according to claim 9, characterized in that The first building block comprises: a fourth determining subunit, configured to determine the gas purchase cost of the park based on the gas purchase price of the park, the output power of the third type of energy conversion equipment, and the conversion efficiency of the first type of energy conversion equipment; wherein the third type of energy conversion equipment refers to equipment that converts gas into electrical energy or thermal energy; a fifth determining subunit, configured to determine the energy storage devices in the energy storage system, and determine the operating cost of the energy storage system based on the number of the energy storage devices and the unit operating costs of the energy storage devices; a sixth determining subunit, configured to determine the electricity purchase cost of the park based on the electricity purchase price of the park and the output power of the external power grid; The fifth construction subunit is used to construct the objective function based on the operating cost of the energy storage system, the electricity purchase cost of the park, the gas purchase cost of the park, the carbon emission coefficient of the external power grid and the carbon emission coefficient of the first type of energy conversion equipment.
14. The device according to claim 9, characterized in that The reward function at least includes: an immediate reward function and a cumulative reward function; the second construction unit includes: a sixth construction subunit, configured to construct the state variable based on the energy load information of the park within the preset time period, the charging power of the energy storage system, the discharging power of the energy storage system, the state of charge information of the energy storage device, the output power prediction value of the wind power system, and the output power prediction value of the photovoltaic system; a seventh constructing subunit, configured to construct the action variable according to the output power increment of the energy storage system and the output power increment of the energy conversion system within the preset time period; The eighth construction subunit is used to construct the reward function of the energy scheduling of the park at any time based on the state variables, the action variables and the objective function to obtain the instant reward function; the ninth construction subunit is used to construct the cumulative reward function based on the instant reward function and preset parameters of the energy scheduling of the park within a preset time period.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes stored computer instructions, wherein when the computer instructions are executed by a processor, the energy scheduling method within the park described in any one of claims 1 to 7 is implemented.
16. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the energy scheduling method within the park as described in any one of claims 1 to 7.
Citation Information
Cited By
Park comprehensive energy low-carbon optimization operation method and system
CN121417334A