A comprehensive energy system anti-fragile optimization operation method and system based on security reinforcement learning, an electronic device, and a medium

By using a security reinforcement learning-based approach, an antifragile optimization model for a comprehensive energy system was established, and a network-based solution algorithm was designed. This solved the system's optimization scheduling problem under uncertainty, improving both economic efficiency and antifragility.

CN119809030BActive Publication Date: 2026-04-14STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2024-12-17
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies suffer from increased uncertainty and insufficient antifragility in the optimal scheduling of integrated energy systems. Traditional methods involve large computational loads or cannot guarantee global optimality, making them difficult to handle complex dynamic systems.

Method used

A security-based reinforcement learning approach is adopted to establish an antifragile optimization model for a comprehensive energy system. By designing an economic value network, a security value network, a policy network, and a target policy network, security reinforcement learning is used to solve the model, thereby improving the system's antifragility.

Benefits of technology

While ensuring security, the system's economy and ability to cope with uncertainty have been improved, achieving better antifragility optimization results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809030B_ABST
    Figure CN119809030B_ABST
Patent Text Reader

Abstract

The application provides a comprehensive energy system anti-fragile optimization operation method and system based on safety reinforcement learning, an electronic device and a medium, and belongs to the technical field of comprehensive energy system operation analysis, and comprises the following steps: step S1, a mathematical model of a comprehensive energy system anti-fragile optimization problem is established, and an objective function and a constraint condition of the mathematical model of the comprehensive energy system anti-fragile optimization problem are set; step S2, the comprehensive energy system anti-fragile optimization problem is converted into an optimization model based on a constraint Markov decision process; step S3, a safety reinforcement learning solving algorithm is designed, each network is updated, and an optimal solution of the comprehensive energy system anti-fragile optimization problem is obtained after convergence. The comprehensive energy system anti-fragile optimization operation method, system, electronic device and medium based on safety reinforcement learning can improve the anti-fragile ability of the system to cope with uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated energy system operation analysis technology, and in particular to an antifragile optimization operation method, system, electronic device and medium for integrated energy systems based on security reinforcement learning. Background Technology

[0002] Energy is crucial to modern society. Integrated energy systems, which combine multiple energy forms with control technologies, are an important way to improve energy security, economic efficiency, and environmental protection. With the development of clean and renewable energy technologies, and the emergence of concepts such as the energy internet and smart grids, the concept and theory of integrated energy systems have been gradually deepened and expanded. In recent years, the high penetration rate of new energy sources and the large-scale integration of power electronic equipment have increased the uncertainty of the optimal scheduling problem of integrated energy systems, making scheduling and control more difficult. Therefore, how to address the uncertainty of integrated energy systems and enhance their anti-fragility capabilities is a key issue that needs to be addressed.

[0003] In recent years, traditional optimization scheduling methods for integrated energy systems have included analytical methods and heuristic algorithms. Analytical methods rely on rigorous mathematical theory to solve problems precisely. However, for large-scale, nonlinear, multivariable, and multi-constraint optimization problems, the computational load can be extremely high or even unsolvable. Furthermore, the solution process often requires simplification and assumptions about the actual problem, leading to a decline in solution quality and significant deviations from reality. Heuristic algorithms do not require complex mathematical models and are relatively simple to implement, but they typically only find approximate or locally optimal solutions, failing to guarantee global optimality. Moreover, these algorithms are overly dependent on parameters, making them unsuitable for large-scale practical applications. With the intelligent and digital development of energy scheduling, artificial intelligence methods, led by reinforcement learning, have been widely applied in the field due to their "model-free" advantage. While reinforcement learning can reduce the impact of model errors and uncertainties on scheduling strategies, it is necessary to enhance the antifragility of increasingly complex dynamic integrated energy systems to achieve safe and economical scheduling. Summary of the Invention

[0004] The purpose of this invention is to provide a method, system, electronic device and medium for antifragile optimization of integrated energy systems based on security reinforcement learning, which can improve the system's antifragility in the face of uncertainty.

[0005] To achieve the above objectives, this invention provides an antifragile optimization operation method for integrated energy systems based on security reinforcement learning, comprising the following steps:

[0006] Step S1: Establish a mathematical model for the antifragile optimization problem of the integrated energy system. Under five constraints—equipment operation constraints, external energy supply constraints, energy storage device constraints, electrical balance constraints, and thermal balance constraints—set the objective function of the mathematical model for the antifragile optimization problem of the integrated energy system.

[0007] Step S2: Transform the antifragile optimization problem of the integrated energy system into an optimization model based on a constrained Markov decision process;

[0008] Step S3: Design a security reinforcement learning algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a goal policy network. Update each network and obtain the optimal solution to the antifragile optimization problem of the integrated energy system after convergence.

[0009] Preferably, in step S1, the objective function is as follows:

[0010]

[0011] F BES (t)=α BES |P BES (t)| (4);

[0012] Where F is the system operating cost; T is the number of time periods in the operating cycle; F NG (t), F grid (t), F BES (t) represents the gas purchase cost, electricity purchase cost, and energy storage depreciation cost at time t, respectively; The unit prices for gas and electricity at time t are respectively; M GB (t), M CHP (t) represents the total amount of natural gas consumed by the gas-fired boiler and the combined heat and power unit at time t; P grid (t) represents the total amount of electricity purchased from the upper-level power grid at time t; α BES P is the depreciation cost factor for energy storage devices. BES (t) represents the charging and discharging power of the energy storage device at time t, with positive values ​​for discharging and negative values ​​for charging.

[0013] Preferably, in step S1, the equipment operation constraints are as follows:

[0014]

[0015] in, These are the minimum and maximum values ​​of the output electrical power of the combined heat and power unit, respectively; P CHP (t) represents the electrical power output of the combined heat and power unit at time t; P CHP (t-1) represents the electrical power output of the cogeneration unit at time t-1; These are the minimum and maximum output thermal power of the electric boiler, respectively; H EB (t) represents the thermal power output of the electric boiler at time t; H EB (t-1) represents the thermal power output of the electric boiler at time t-1; These are the minimum and maximum output thermal power of the gas-fired boiler, respectively; H GB (t) represents the thermal power output of the gas-fired boiler at time t; H GB (t-1) represents the thermal power output of the gas-fired boiler at time t-1; These are the minimum and maximum values ​​of the output power of the energy storage device, respectively; P BES (t) represents the electrical power output by the energy storage device at time t; These are the upper and lower limits of ramp-up constraints for combined heat and power units; These are the upper and lower limits of the ramp-up constraint for electric boilers; These are the upper and lower limits of the ramp-up constraint for gas-fired boilers.

[0016] Preferably, in step S1, the external energy supply constraints are as follows:

[0017]

[0018] in, These represent the minimum and maximum values ​​of the power exchanged between the integrated energy system and the upper-level power grid, respectively; P grid (t) represents the exchange power between the integrated energy system and the upper-level power grid at time t.

[0019] Preferably, in step S1, the constraints of the energy storage device are as follows:

[0020]

[0021] Among them, C SOC (t) represents the state of charge of the energy storage device at time t; C SOC (t-1) represents the state of charge of the energy storage device at time t-1; η BES Q is the charge / discharge coefficient; BES η represents the capacity of the energy storage device; Δt represents the time interval between adjacent dispatch periods; η represents the capacity of the energy storage device. ch η is the charging coefficient; dis The discharge coefficient; These are the minimum and maximum charge values, respectively.

[0022] Preferably, in step S1, the power balance constraints are as follows:

[0023] P grid (t)+P CHP (t)+P PV (t)+P BES(t)-P EB (t)=P load (t)(16);

[0024] Among them, P PV (t) represents the output power of the photovoltaic unit at time t; P load (t) represents the electrical load demand that the integrated energy system needs to meet at time t; P EB (t) represents the electrical power consumed by the electric boiler at time t; P grid (t) represents the electrical power supplied by the upstream power grid at time t.

[0025] Preferably, in step S1, the thermal energy balance constraint is as follows:

[0026] H CHP (t)+H GB (t)+H EB (t)=H load (t)(17);

[0027] Among them, H CHP (t) represents the output thermal power of the cogeneration unit at time t; H load (t) represents the heat load demand that the integrated energy system needs to meet at time t.

[0028] Preferably, in step S2, the optimization model based on the constrained Markov decision process is expressed as follows:

[0029]

[0030] Where Q(π) is the future reward discount cost; E π Let r(s) be the expected function under strategy π; γ is the discount factor; t ,a t ) is an instant reward; s t The state at time t; a t Let be the action at time t; C(π) be the future discounted safety cost; and d be the constraint tolerance.

[0031] Preferably, in step S2, the constrained Markov decision process includes a state space, an action space, a reward function, a state transition design, a discount factor, and a safety constraint function.

[0032] Preferably, the state space includes the electrical load demand, heat load demand, output power of the photovoltaic unit, and state of charge of the energy storage device at time t-1, which the integrated energy system needs to meet at time t. Random disturbances are superimposed on historical source-load data to simulate uncertainties and enhance the system's anti-fragility. Specifically, it is expressed as follows:

[0033] s t =[P load(t),H load (t),P PV (t),C SOC (t-1)] (19).

[0034] Preferably, the operating space includes the electrical power output of the combined heat and power unit, the thermal power output of the gas boiler, and the electrical power output of the electric energy storage, specifically expressed as follows:

[0035] a t =[P CHP (t),H GB (t),P BES (t)] (20).

[0036] Preferably, the negative of the operating cost of the integrated energy system at time t is taken as the reward function for reinforcement learning, specifically expressed as:

[0037] r(s t ,a t )=-ξ(F NG (t)+F grid (t)+F BES (t)) (21);

[0038] Where ξ is the scaling factor for instant rewards;

[0039] Define the future reward discount cost Q(π) as follows:

[0040]

[0041] Preferably, the security constraint function is as follows:

[0042]

[0043] Where ξ1 is the safety-limited scaling factor; operator |a + =max(0,a); P x (t) represents the active power output of unit x at time t; These represent the minimum and maximum output values ​​of unit x at any given time.

[0044] Preferably, the update process for the economic value network, security value network, policy network, economic target network, security target network, and target policy network is as follows:

[0045] The economic value network update process is as follows:

[0046]

[0047] Wherein, L(φ) Q ) represents the economic value network loss function; y tLet Q(s) be the target economic value at time t; t ,α t |φ Q φ is the action value function output by the economic value network; Q These are the economic value network parameters before this update; The parameters of the economic value network after this round of updates; α φ ∠ is the learning rate of the economic value network; ∠ is the gradient calculation function;

[0048] The security value network update process is as follows:

[0049]

[0050] Wherein, L(ψ) C ) represents the loss function of the security value network; C(s) represents the target safety value at time t. t ,a t ∣ψ C ψ represents the safety value function under different constraints; C These are the security value network parameters prior to this update. The parameters for the security value network after this update; α ψ For security value network learning rate;

[0051] The policy network update process is as follows:

[0052]

[0053] Wherein, L(s) t ,a t |λ) is the policy network loss function; λ is the Lagrange multiplier; θ π These are the policy network parameters before this round of updates; These are the policy network parameters after this round of updates; π The learning rate of the policy network;

[0054] The update process for the economic target network, security target network, and target policy network is as follows:

[0055]

[0056] in, Here are the network parameters for the economic objectives after this round of updates; τ is the soft update coefficient. These are the network parameters for the economic targets prior to this update. These are the network parameters for the security target after this round of updates; These are the network parameters for the security target prior to this update. These are the network parameters for the target policy after this round of updates; These are the network parameters for the target policy prior to this update.

[0057] After each network update converges, the optimal solution to the antifragile optimization problem of the integrated energy system is obtained.

[0058] Preferably, the integrated energy system includes an energy input unit, an energy conversion unit, an energy storage unit, and an energy output unit. The energy input unit includes the upstream power grid, the external gas grid, and photovoltaic power generation devices. The energy conversion equipment includes combined heat and power units, electric boilers, and gas boilers. The energy storage unit includes electric energy storage devices. The energy output unit includes electrical load and heat load.

[0059] This invention also provides an antifragile optimization operation system for integrated energy systems based on security reinforcement learning, comprising:

[0060] The modeling module is used to establish a mathematical model of the antifragile optimization problem of the integrated energy system. Under five constraints, namely equipment operation constraints, external energy supply constraints, electric energy storage device constraints, electrical balance constraints, and thermal balance constraints, the objective function of the mathematical model of the antifragile optimization problem of the integrated energy system is set.

[0061] The transformation module transforms the antifragile optimization problem of integrated energy systems into an optimization model based on constrained Markov decision processes.

[0062] The solution module is used to design a secure reinforcement learning solution algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a goal policy network. It updates each network and obtains the optimal solution to the antifragile optimization problem of the integrated energy system after convergence.

[0063] The present invention also provides a computer device, including: a memory and a processor; the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for antifragile optimization of integrated energy systems based on security reinforcement learning.

[0064] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method for antifragile optimization of integrated energy systems based on security reinforcement learning.

[0065] Therefore, the present invention employs the above-mentioned method, system, electronic device, and medium for antifragile optimization of integrated energy systems based on security reinforcement learning, and the beneficial technical effects are as follows:

[0066] (1) This invention designs a safe reinforcement learning method for solving the optimization scheduling problem of integrated energy system, which can improve the economy of the system while ensuring safety;

[0067] (2) In the process of establishing a mathematical model, the present invention superimposes randomness on historical data, so that it can flexibly cope with the uncertainty fluctuations of the system and has a good antifragile optimization capability. Attached Figure Description

[0068] Figure 1 This is a power supply architecture diagram of the test system in Example 1;

[0069] Figure 2 This is a graph showing the convergence of the reward function. Detailed Implementation

[0070] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0071] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0072] Example 1

[0073] An antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning includes the following steps:

[0074] Step S1: Establish a mathematical model for the antifragile optimization problem of the integrated energy system. Under five constraints—equipment operation constraints, external energy supply constraints, electric energy storage device constraints, electrical balance constraints, and thermal balance constraints—set the objective function of the mathematical model for the antifragile optimization problem of the integrated energy system.

[0075] The objective function is as follows:

[0076]

[0077] F BES (t)=α BES |P BES (t)| (4);

[0078] Where F is the system operating cost; T is the number of time periods in the operating cycle; F NG (t), F grid (t), F BES (t) represents the gas purchase cost, electricity purchase cost, and energy storage depreciation cost at time t, respectively; The unit prices for gas and electricity at time t are respectively; M GB (t), M CHP (t) represents the total amount of natural gas consumed by the gas-fired boiler and the combined heat and power unit at time t; P grid (t) represents the total amount of electricity purchased from the upper-level power grid at time t; α BES P is the depreciation cost factor for energy storage devices. BES(t) represents the charging and discharging power of the energy storage device at time t, with positive values ​​for discharging and negative values ​​for charging.

[0079] The equipment operating constraints are as follows:

[0080]

[0081]

[0082] in, These are the minimum and maximum values ​​of the output electrical power of the combined heat and power unit, respectively; P CHP (t) represents the electrical power output of the combined heat and power unit at time t; P CHP (t-1) represents the electrical power output of the cogeneration unit at time t-1; These are the minimum and maximum output thermal power of the electric boiler, respectively; H EB (t) represents the thermal power output of the electric boiler at time t; H EB (t-1) represents the thermal power output of the electric boiler at time t-1; These are the minimum and maximum output thermal power of the gas-fired boiler, respectively; H GB (t) represents the thermal power output of the gas-fired boiler at time t; H GB (t-1) represents the thermal power output of the gas-fired boiler at time t-1; These are the minimum and maximum values ​​of the output power of the energy storage device, respectively; P BES (t) represents the electrical power output by the energy storage device at time t; These are the upper and lower limits of ramp-up constraints for combined heat and power units; These are the upper and lower limits of the ramp-up constraint for electric boilers; These are the upper and lower limits of the ramp-up constraint for gas-fired boilers.

[0083] External energy supply constraints are as follows:

[0084]

[0085] in, These represent the minimum and maximum values ​​of the power exchanged between the integrated energy system and the upper-level power grid, respectively; P grid (t) represents the exchange power between the integrated energy system and the upper-level power grid at time t.

[0086] The constraints of energy storage devices are as follows:

[0087]

[0088] Among them, C SOC (t) represents the state of charge of the energy storage device at time t; C SOC(t-1) represents the state of charge of the energy storage device at time t-1; η BES Q is the charge / discharge coefficient; BES η represents the capacity of the energy storage device; Δt represents the time interval between adjacent dispatch periods; η represents the capacity of the energy storage device. ch η is the charging coefficient; dis The discharge coefficient; These are the minimum and maximum charge values, respectively.

[0089] The energy balance constraints are as follows:

[0090] P grid (t)+P CHP (t)+P PV (t)+P BES (t)-P EB (t)=P load (t)(16);

[0091] Among them, P PV (t) represents the output power of the photovoltaic unit at time t; P load (t) represents the electrical load demand that the integrated energy system needs to meet at time t; P EB (t) represents the electrical power consumed by the electric boiler at time t; P grid (t) represents the electrical power supplied by the upstream power grid at time t.

[0092] In step S1, the thermal energy balance constraints are as follows:

[0093] H CHP (t)+H GB (t)+H EB (t)=H load (t)(17);

[0094] Among them, H CHP (t) represents the output thermal power of the cogeneration unit at time t; H load (t) represents the heat load demand that the integrated energy system needs to meet at time t.

[0095] Step S2: Transform the antifragile optimization problem of the integrated energy system into an optimization model based on a constrained Markov decision process.

[0096] The optimization model based on the constrained Markov decision process is expressed as follows:

[0097]

[0098] Where Q(π) is the future reward discount cost; E π Let r(s) be the expected function under strategy π; γ is the discount factor; t ,a t ) is an instant reward; st The state at time t; a t Let be the action at time t; C(π) be the future discounted safety cost; and d be the constraint tolerance.

[0099] Constrained Markov decision processes include state space, action space, reward function, state transition design, discount factor, and safety constraint function.

[0100] The state space contains the state information needed by the agent to decide its next action. The state space includes the electrical load demand, heat load demand, photovoltaic power output of the integrated energy system at time t, and the state of charge of the energy storage device at time t-1. Random disturbances are superimposed on historical source-load data to simulate uncertainty and enhance the system's anti-fragility. Specifically, it is represented as follows:

[0101] s t =[P load (t),H load (t),P PV (t),C SOC (t-1)](19).

[0102] The action space is the set of all actions that the agent can perform in the current state, which directly affects the agent's training process. The action space includes the electrical power output of the combined heat and power (CHP) unit, the thermal power output of the gas-fired boiler, and the electrical power output of the electric energy storage unit, specifically represented as follows:

[0103] a t =[P CHP (t),H GB (t),P BES (t)](20).

[0104] The goal of reinforcement learning is to maximize the reward function. Therefore, the negative of the system operating cost at time t is taken as the reward function for reinforcement learning, specifically expressed as:

[0105] r(s t ,a t )=-ξ(F NG (t)+F grid (t)+F BES (t))(21);

[0106] Where ξ is the scaling factor for instant rewards;

[0107] Based on this, the future reward discount cost Q(π) is defined as:

[0108]

[0109] State transition design:

[0110] A state transition function describes the probability that an agent will transition to the next state after taking an action, and it is usually related to the uncertainty of the environment.

[0111] Discount factor:

[0112] The discount factor is used to calculate the current value of future rewards, enabling agents to balance short-term and long-term interests.

[0113] The security restriction functions are as follows:

[0114]

[0115] Where ξ1 is the safety-limited scaling factor; operator |a| + =max(0,a); P x (t) represents the active power output of unit x at time t; These represent the minimum and maximum output values ​​of unit x at any given time.

[0116] Step S3: Design a security reinforcement learning algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a goal policy network. Update each network and obtain the optimal solution to the antifragile optimization problem of the integrated energy system after convergence.

[0117] The economic value network update process is as follows:

[0118]

[0119] Wherein, L(φ) Q ) represents the economic value network loss function; y t Let Q(s) be the target economic value at time t; t ,α t |φ Q φ is the action value function output by the economic value network; Q These are the economic value network parameters before this update; The parameters of the economic value network after this round of updates; α φ ∠ is the learning rate of the economic value network; ∠ is the gradient calculation function.

[0120] The security value network update process is as follows:

[0121]

[0122] Wherein, L(ψ) C ) represents the loss function of the security value network; C(s) represents the target safety value at time t. t ,a t ∣ψ Cψ represents the safety value function under different constraints; C These are the security value network parameters prior to this update. The parameters for the security value network after this update; α ψ The learning rate of the network is for security value.

[0123] The policy network update process is as follows:

[0124]

[0125] Wherein, L(s) t ,a t |λ) is the policy network loss function; λ is the Lagrange multiplier; θ π These are the policy network parameters before this round of updates; These are the policy network parameters after this round of updates; π is the learning rate of the policy network.

[0126] The update process for the economic target network, security target network, and target policy network is as follows:

[0127]

[0128] in, Here are the network parameters for the economic objectives after this round of updates; τ is the soft update coefficient. These are the network parameters for the economic targets prior to this update. These are the network parameters for the security target after this round of updates; These are the network parameters for the security target prior to this update. These are the network parameters for the target policy after this round of updates; These are the network parameters for the target strategy prior to this update.

[0129] After each network update converges, the optimal solution to the antifragile optimization problem of the integrated energy system is obtained.

[0130] This invention applies an antifragile optimization operation method for integrated energy systems based on security reinforcement learning to... Figure 1 The method is further illustrated in the test system shown. Figure 1The test system shown is an electrothermal coupled integrated energy system. This system includes an energy input unit, an energy conversion unit, an energy storage unit, and an energy output unit. The energy input unit includes an upstream power grid, an external gas grid, and photovoltaic power generation devices. The upstream power grid provides electricity, the external gas grid provides natural gas, and the photovoltaic power generation devices provide electricity converted from solar energy. The energy conversion equipment includes a combined heat and power (CHP) unit, an electric boiler, and a gas-fired boiler. Natural gas is converted into electricity and heat (CHP unit), electricity is converted into heat (electric boiler), and natural gas is converted into heat (gas-fired boiler). The energy storage unit includes electrical energy storage devices, and the energy output unit includes electrical load and heat load. The source-load data is actual historical data from an industrial park in China.

[0131] Based on historical data and the internal structure of the system, a constrained Markov decision process is constructed.

[0132] Based on the constrained Markov decision process, the model described in this invention is solved using a secure reinforcement learning algorithm to obtain the convergence of the algorithm's reward function, as follows: Figure 2 As shown.

[0133] Example 2

[0134] An antifragile optimization operation system for integrated energy systems based on security reinforcement learning, comprising:

[0135] The modeling module is used to establish a mathematical model of the antifragile optimization problem of the integrated energy system. Under five constraints, namely equipment operation constraints, external energy supply constraints, electric energy storage device constraints, electrical balance constraints, and thermal balance constraints, the objective function of the mathematical model of the antifragile optimization problem of the integrated energy system is set.

[0136] The transformation module transforms the antifragile optimization problem of integrated energy systems into an optimization model based on constrained Markov decision processes.

[0137] The solution module is used to design a secure reinforcement learning solution algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a goal policy network. It updates each network and obtains the optimal solution to the antifragile optimization problem of the integrated energy system after convergence.

[0138] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0140] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0141] It is worth noting that all contents not described in detail in this invention are existing technologies and are well known to those skilled in the art.

[0142] Therefore, the present invention employs the above-mentioned method, system, electronic device and medium for antifragile optimization of integrated energy system based on security reinforcement learning, which can improve the system's antifragility in the face of uncertainty.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for antifragile optimization operation of a comprehensive energy system based on security reinforcement learning, characterized in that, Includes the following steps: Step S1: Establish a mathematical model for the antifragile optimization problem of the integrated energy system. Under five constraints—equipment operation constraints, external energy supply constraints, energy storage device constraints, electrical balance constraints, and thermal balance constraints—set the objective function of the mathematical model for the antifragile optimization problem of the integrated energy system. Step S2: Transform the antifragile optimization problem of the integrated energy system into an optimization model based on a constrained Markov decision process; Step S3: Design a security reinforcement learning algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a target policy network. Update each network and obtain the optimal solution to the antifragile optimization problem of the integrated energy system after convergence. The update process for the economic value network, security value network, policy network, economic goal network, security goal network, and goal policy network is as follows: The economic value network update process is as follows: L(φ Q )=[y t -Q(s t ,a t ∣φ Q )] 2 (24); Wherein, L(φ) Q ) represents the economic value network loss function; y t Let Q(s) be the target economic value at time t; t ,α t |φ Q φ is the action value function output by the economic value network; Q These are the economic value network parameters before this update; The parameters of the economic value network after this round of updates; α φ ∠ is the learning rate of the economic value network; ∠ is the gradient calculation function; The security value network update process is as follows: Wherein, L(ψ) C ) represents the loss function of the security value network; C(s) represents the target safety value at time t. t ,a t ∣ψ C ψ represents the safety value function under different constraints; C These are the security value network parameters prior to this update. The parameters for the security value network after this update; α ψ For security value network learning rate; The policy network update process is as follows: L(s t ,a t ∣λ)=Q(s t ,a t ∣φ Q )-λ[C(s t ,a t ∣ψ C )-d] (28); Wherein, L(s) t ,a t |λ) is the policy network loss function; λ is the Lagrange multiplier; θ π These are the policy network parameters before this round of updates; These are the policy network parameters after this round of updates; π Let d be the learning rate of the policy network, and d be the constraint tolerance. The update process for the economic target network, security target network, and target policy network is as follows: in, Here are the network parameters for the economic objectives after this round of updates; τ is the soft update coefficient. These are the network parameters for the economic targets prior to this update. These are the network parameters for the security target after this round of updates; These are the network parameters for the security target prior to this update. These are the network parameters for the target policy after this round of updates; These are the network parameters for the target policy prior to this update.

2. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 1, characterized in that, In step S1, the objective function is as follows: F BES (t)=α BES |P BES (t)| (4); Where F is the system operating cost; T is the number of time periods in the operating cycle; F NG (t), F grid (t), F BES (t) represents the gas purchase cost, electricity purchase cost, and energy storage depreciation cost at time t, respectively; The unit prices for gas and electricity at time t are respectively; M GB (t), M CHP (t) represents the total amount of natural gas consumed by the gas-fired boiler and the combined heat and power unit at time t; P grid (t) represents the total amount of electricity purchased from the upper-level power grid at time t; α BES P is the depreciation cost factor for energy storage devices. BES (t) represents the charging and discharging power of the energy storage device at time t, with positive values ​​for discharging and negative values ​​for charging.

3. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 2, characterized in that, In step S1, the equipment operation constraints are as follows: in, These are the minimum and maximum values ​​of the output electrical power of the combined heat and power unit, respectively; P CHP (t) represents the electrical power output of the combined heat and power unit at time t; P CHP (t-1) represents the electrical power output of the cogeneration unit at time t-1; These are the minimum and maximum output thermal power of the electric boiler, respectively; H EB (t) represents the thermal power output of the electric boiler at time t; H EB (t-1) represents the thermal power output of the electric boiler at time t-1; These are the minimum and maximum output thermal power of the gas-fired boiler, respectively; H GB (t) represents the thermal power output of the gas-fired boiler at time t; H GB (t-1) represents the thermal power output of the gas-fired boiler at time t-1; These are the minimum and maximum values ​​of the output power of the energy storage device, respectively; P BES (t) represents the electrical power output by the energy storage device at time t; These are the upper and lower limits of ramp-up constraints for combined heat and power units; These are the upper and lower limits of the ramp-up constraint for electric boilers; These are the upper and lower limits of the ramp-up constraint for gas-fired boilers.

4. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 3, characterized in that, In step S1, the external energy supply constraints are as follows: in, These represent the minimum and maximum values ​​of the power exchanged between the integrated energy system and the upper-level power grid, respectively; P grid (t) represents the exchange power between the integrated energy system and the upper-level power grid at time t.

5. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 4, characterized in that, In step S1, the constraints on the energy storage device are as follows: Among them, C SOC (t) represents the state of charge of the energy storage device at time t; C SOC (t-1) represents the state of charge of the energy storage device at time t-1; η BES Q is the charge / discharge coefficient; BES η represents the capacity of the energy storage device; Δt represents the time interval between adjacent dispatch periods; η represents the capacity of the energy storage device. ch η is the charging coefficient; dis The discharge coefficient; These are the minimum and maximum charge values, respectively.

6. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 5, characterized in that, In step S1, the energy balance constraints are as follows: P grid (t)+P CHP (t)+P PV (t)+P BES (t)-P EB (t)=P load (t) (16); Among them, P PV (t) represents the output power of the photovoltaic unit at time t; P load (t) represents the electrical load demand that the integrated energy system needs to meet at time t; P EB (t) represents the electrical power consumed by the electric boiler at time t; P grid (t) represents the electrical power supplied by the upstream power grid at time t.

7. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 6, characterized in that, In step S1, the thermal energy balance constraints are as follows: H CHP (t)+H GB (t)+H EB (t)=H load (t) (17); Among them, H CHP (t) represents the output thermal power of the cogeneration unit at time t; H load (t) represents the heat load demand that the integrated energy system needs to meet at time t.

8. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 7, characterized in that, In step S2, the optimization model based on the constrained Markov decision process is expressed as follows: Where Q(π) is the future reward discount cost; E π Let r(s) be the expected function under strategy π; γ is the discount factor; t ,a t ) is an instant reward; s t The state at time t; a t Let be the action at time t; C(π) be the future discounted safety cost; and d be the constraint tolerance.

9. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 8, characterized in that, In step S2, the constrained Markov decision process includes the state space, action space, reward function, state transition design, discount factor, and safety constraint function.

10. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 9, characterized in that, The state space includes the electrical load demand, heat load demand, photovoltaic power output, and state of charge of the energy storage device at time t-1, which the integrated energy system needs to meet. Random disturbances are superimposed on historical source-load data to simulate uncertainties and enhance the system's anti-fragility. Specifically, it is represented as follows: s t =[P load (t),H load (t),P PV (t),C SOC (t-1)] (19)。 11. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 10, characterized in that, The operating space includes the electrical power output of the combined heat and power unit, the thermal power output of the gas-fired boiler, and the electrical power output of the electric energy storage, specifically expressed as: a t =[P CHP (t),H GB (t),P BES (t)] (20)。 12. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 11, characterized in that, The negative of the overall energy system operating cost at time t is taken as the reward function for reinforcement learning, specifically expressed as: r(s t ,a t )=-ξ(F NG (t)+F grid (t)+F BES (t))(21); Where ξ is the scaling factor for instant rewards; Define the future reward discount cost Q(π) as follows:

13. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 12, characterized in that, The security restriction functions are as follows: Where ξ1 is the safety-limited scaling factor; operator |a| + =max(0,a); P x (t) represents the active power output of unit x at time t; These represent the minimum and maximum output values ​​of unit x at any given time.

14. The antifragile optimization operation method for a comprehensive energy system based on security reinforcement learning according to claim 13, characterized in that, An integrated energy system includes an energy input unit, an energy conversion unit, an energy storage unit, and an energy output unit. The energy input unit includes the upstream power grid, the external gas grid, and photovoltaic power generation devices. The energy conversion equipment includes combined heat and power units, electric boilers, and gas boilers. The energy storage unit includes electric energy storage equipment. The energy output unit includes electrical load and heat load.

15. A comprehensive energy system antifragile optimization operation system based on security reinforcement learning, characterized in that, A method for implementing an antifragile optimization operation of a comprehensive energy system based on security reinforcement learning as described in any one of claims 1-14, comprising: The modeling module is used to establish a mathematical model of the antifragile optimization problem of the integrated energy system. Under five constraints, namely equipment operation constraints, external energy supply constraints, electric energy storage device constraints, electrical balance constraints, and thermal balance constraints, the objective function of the mathematical model of the antifragile optimization problem of the integrated energy system is set. The transformation module transforms the antifragile optimization problem of integrated energy systems into an optimization model based on constrained Markov decision processes. The solution module is used to design a secure reinforcement learning solution algorithm, including an economic value network, an economic goal network, a security value network, a security goal network, a policy network, and a goal policy network. It updates each network and obtains the optimal solution to the antifragile optimization problem of the integrated energy system after convergence.

16. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that when the processor executes the computer program, it implements the steps of the antifragile optimization operation method for an integrated energy system based on security reinforcement learning as described in any one of claims 1-14.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When a computer program is executed by a processor, it implements the steps of the antifragile optimization operation method for an integrated energy system based on security reinforcement learning as described in any one of claims 1-14.