Integrated energy system double-side collaborative control method, device and storage medium
Patent Information
- Application Number
- CN202311192263.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-09-15
AI Technical Summary
[0017]By collecting real-time status data from the integrated energy system, this data is input into a trained deep policy model. Through the policy network, decision-making data, including equipment operating parameters on the energy supply side and controllable power consumption on the load side, are obtained. The operation of each device on the energy supply side is controlled based on the equipment operating parameters, and the power consumption on the load side and the internal temperature of the heating area are controlled based on the controllable power consumption. This achieves coordinated optimization control of the load and energy supply sides of the integrated energy system. By dividing the load side electrical load into controllable and uncontrollable loads, and scheduling the controllable load, the system can further reduce operating costs while ensuring responsiveness to load demand and meeting user thermal comfort. Furthermore, this method improves the deep deterministic policy gradient algorithm through a priority experience replay strategy and L2 regularization strategy, thereby training a deep policy model that further improves the accuracy of the model output and ensures the reliability of the two-sided coordinated control.
Smart Images

Figure CN117134432B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated energy system control technology, and in particular to a dual-sided collaborative control method, device and storage medium for integrated energy systems. Background Technology
[0002] Integrated energy systems with multi-energy coordination and complementarity capabilities are an important technological path to achieving "dual carbon" goals. An integrated energy system refers to a multi-energy integrated utilization system formed by coordinating and optimizing the production, transmission, conversion, and consumption of various types of energy, such as electricity, gas, and heat, during the planning, construction, and operation phases. It consists of energy supply units, conversion units, storage units, and multiple loads. The coexistence and coupled use of multiple energy units in the system have made the optimization and control of integrated energy systems, thereby achieving optimal system operation, a matter of great concern.
[0003] However, the problem of integrated energy system optimization and control faces a variety of uncertainties. The randomness of user load and the volatility of renewable energy output can have a significant impact on system operation optimization, making it more difficult to balance constraints and carry out multi-energy optimization and control.
[0004] Traditional control methods, such as robust optimization, stochastic programming, and model predictive control, rely on the distribution knowledge or predictive information of these uncertainties. Since obtaining accurate distribution knowledge or predictive information is extremely difficult and impractical, control strategies are limited by the accuracy of this knowledge or prediction, and cannot accurately respond to dynamic changes in the actual environment. Furthermore, to achieve coordinated operation between the supply and demand sides of an integrated energy system, it is necessary to study collaborative control strategies for both sides.
[0005] In view of this, the present invention is hereby proposed. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method, device, and storage medium for dual-side coordinated control of an integrated energy system. By coordinating the control of energy supply-side resources and load-side adjustable resources, it fully leverages the adjustment potential of load-side dispatchable resources, addresses the uncertainties of renewable energy generation on the energy supply side, uncontrollable electrical load demand on the load side, and external environmental temperature in the integrated energy system, and achieves dual-side coordinated control of the integrated energy system. This not only ensures responsiveness to load-side demand and meets user thermal comfort requirements but also further reduces system operating costs and guarantees the reliability of dual-side coordinated control.
[0007] This invention provides a dual-sided coordinated control method for an integrated energy system, the method comprising:
[0008] S1. Collect real-time status data of the integrated energy system. The real-time status data includes uncontrollable electrical load demand, renewable energy power generation, state of charge of the energy storage device, electricity purchase price, gas purchase price, internal temperature of the heating area, and external ambient temperature.
[0009] S2. Input the real-time status data into a pre-trained deep policy model, and determine the decision action data of the integrated energy system through the policy network in the deep policy model. The decision action data includes the equipment operating parameters on the energy supply side and the controllable power consumption on the load side of the integrated energy system.
[0010] S3. Control the operation of each device on the power supply side based on the device operating parameters, and control the power consumed on the load side of the integrated energy system based on the controllable power consumption.
[0011] The deep policy model is trained based on an improved deep deterministic policy gradient algorithm, which combines a priority experience replay strategy and an L2 regularization strategy. The priority experience replay strategy is used to assign a corresponding priority to each sample in the experience replay pool, and the L2 regularization strategy is used to add penalty terms to the performance objective function and loss function to optimize the parameters of the deep policy model.
[0012] This invention provides an electronic device, the electronic device comprising:
[0013] Processor and memory;
[0014] The processor executes the steps of the integrated energy system dual-side cooperative control method described in any embodiment by calling the program or instructions stored in the memory.
[0015] This invention provides a computer-readable storage medium storing a program or instructions that cause a computer to execute the steps of the integrated energy system dual-side coordinated control method described in any embodiment.
[0016] The embodiments of the present invention have the following technical effects:
[0017] By collecting real-time status data from the integrated energy system, this data is input into a trained deep policy model. Through the policy network, decision-making data, including equipment operating parameters on the energy supply side and controllable power consumption on the load side, are obtained. The operation of each device on the energy supply side is controlled based on the equipment operating parameters, and the power consumption on the load side and the internal temperature of the heating area are controlled based on the controllable power consumption. This achieves coordinated optimization control of the load and energy supply sides of the integrated energy system. By dividing the load side electrical load into controllable and uncontrollable loads, and scheduling the controllable load, the system can further reduce operating costs while ensuring responsiveness to load demand and meeting user thermal comfort. Furthermore, this method improves the deep deterministic policy gradient algorithm through a priority experience replay strategy and L2 regularization strategy, thereby training a deep policy model that further improves the accuracy of the model output and ensures the reliability of the two-sided coordinated control. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A schematic diagram of an integrated energy system provided in an embodiment of the present invention;
[0020] Figure 2 This is a flowchart of a dual-sided collaborative control method for an integrated energy system provided in an embodiment of the present invention;
[0021] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0023] The integrated energy system dual-side coordinated control method provided in this invention is mainly applicable to situations where the load side and energy supply side of an integrated energy system need to be controlled. This integrated energy system dual-side coordinated control method can be executed by an electronic device integrated into a computer, smartphone, or server.
[0024] Before providing a detailed description of the integrated energy system dual-sided coordinated control method provided in the embodiments of the present invention, an exemplary description of the integrated energy system to which this method is applicable will be given first. Figure 1 This is a schematic diagram of an integrated energy system provided for an embodiment of the present invention.
[0025] like Figure 1 As shown, an integrated energy system can consist of renewable energy power generation equipment, combined heat and power (CHP) units, electric boilers, gas boilers, electric energy storage, and multi-energy loads (electric loads and thermal loads) on the load side. Among these, the electrical load includes uncontrollable electrical loads (i.e., critical loads) and controllable electrical loads, while the thermal load can be understood as the heating load. Because the human body cannot easily perceive small-scale temperature changes, the load side (such as...) Figure 1 The heat load of buildings (in the context of the building) has the flexibility to be scheduled, and the heat load can participate in the system's demand response as an adjustable load.
[0026] In this embodiment of the invention, the load side is no longer studied using a single device (such as an air conditioner or washing machine) within a single end user. Instead, the controllable electrical load and controllable thermal load on the load side of the system are studied as a whole. The park operator can purchase electricity and natural gas from external energy networks or sell surplus electricity to the external power grid. The operator can provide corresponding cash compensation to the loads participating in regulation. The goal of the park operator is to collaboratively manage the dispatchable resources on both the energy supply and load sides, achieving overall economic operation of the system through coordinated efforts.
[0027] Figure 2 This is a flowchart of a dual-sided coordinated control method for an integrated energy system provided in an embodiment of the present invention. See also... Figure 2 The specific methods for dual-side coordinated control of this integrated energy system include:
[0028] S1. Collect real-time status data of the integrated energy system. The real-time status data includes uncontrollable electrical load demand, renewable energy power generation, state of charge of energy storage devices, electricity purchase price, gas purchase price, internal temperature of heating area and external ambient temperature.
[0029] The integrated energy system can consist of renewable energy power generation equipment, combined heat and power units, electric boilers, gas boilers, electric energy storage devices, and multi-energy loads (electric loads and heat loads) in buildings. It can be configured to meet the specific electricity and heat load requirements of industrial parks, residential communities, or factories.
[0030] In this embodiment of the disclosure, the electrical load may include uncontrollable electrical loads (which can be understood as critical loads) and controllable electrical loads, and the heat load is the space heating load. Considering that users are not easily aware of small-scale temperature changes, and because the building heat load has the flexibility to be scheduled, the heat load can participate in the system's demand response as an adjustable load.
[0031] Therefore, in this embodiment, the load side (i.e., the consumer side) no longer manages individual devices (such as air conditioners, washing machines, etc.) of a single end user, but rather studies the controllable electrical load and controllable heat load on the system's energy load side as a whole. The park operator can purchase electricity and natural gas from external energy networks, or sell surplus electricity to the external power grid. The operator will provide corresponding compensation for loads participating in regulation, such as offering discounts to users participating in the regulation of electrical loads. The park operator's goal is to collaboratively manage the dispatchable resources on both the supply and load sides, achieving overall economic operation of the system through coordinated efforts from both sides.
[0032] Specifically, it is necessary to use the real-time status data of the integrated energy system at the current moment to determine the operating parameters of the equipment on the energy supply side and the controllable power consumption on the load side of the integrated energy system for a future period of time. The real-time status data can include uncontrollable electrical load demand, renewable energy generation capacity, state of charge of energy storage devices, electricity purchase price, gas purchase price, internal temperature of the heating area, and external ambient temperature.
[0033] The equipment operating parameters include the output electrical power of the combined heat and power (CHP) unit, the output thermal power of the CHP unit, the output thermal power of the gas boiler, the output thermal power of the electric boiler, the operating power of the electric energy storage device, and the interaction power between the system and the external power grid. Controllable power consumption refers to the actual power consumed by the adjustable electrical load, i.e., the power of the electrical load participating in the regulation.
[0034] In this embodiment of the disclosure, an equipment model and a load-side heat load model of the integrated energy system can be constructed first. The input and output parameters of the deep strategy model are determined through the model, that is, which parameters are included in the real-time status data and which parameters are included in the equipment operating parameters.
[0035] The equipment model of the integrated energy system may include a combined heat and power unit model, a gas boiler model, an electric boiler model, and an electric energy storage device model; the load-side heat load model may be a model of the load-side (i.e., user-side) heat load in the integrated energy system.
[0036] For example, the specific model of a combined heat and power (CHP) unit is as follows:
[0037]
[0038] In the formula, P t CHP η represents the thermal power and electrical power output of the combined heat and power unit during time period t. CHPq η CHPp These are the electrical efficiency and thermal efficiency of the combined heat and power unit, respectively.
[0039] The gas-fired boiler model is as follows:
[0040]
[0041] In the formula, η represents the thermal power output and natural gas consumption of the gas-fired boiler during time period t. GB For the efficiency of gas-fired boilers;
[0042] The electric boiler model is as follows:
[0043]
[0044] In the formula, P t EB η represents the thermal power output and electrical power input of the electric boiler during time period t, respectively. EB For the efficiency of electric boilers;
[0045] The specific model of the energy storage device is as follows:
[0046]
[0047] In the formula, These are the state of charge (SOC) of the energy storage device at time t+1 and time t, respectively, P t BES P is the charging or discharging efficiency of an energy storage device during time period t. t BES A positive value represents the energy storage device discharging, and a negative value represents the energy storage device charging. D BES η is the capacity of the energy storage device, Δt is the time length between two adjacent time periods, i.e., the time interval. BES It is the charging or discharging coefficient of an energy storage device, specifically expressed as:
[0048]
[0049] In the formula, η ch η dis These represent the charging efficiency and discharging efficiency of the energy storage device, respectively.
[0050] An exemplary model of the adjustable electrical load on the load side of an integrated energy system is specifically represented as follows:
[0051]
[0052] In the formula, P t DL To determine the actual electrical power consumed by the adjustable electrical load during time period t. These are the minimum and maximum electrical power that the adjustable electrical load can consume, respectively.
[0053] The load-side heat load model is specifically represented as follows:
[0054]
[0055] In the formula, Let T be the heat load (i.e., heating load) at time t. t in T t out These represent the internal temperature of the heating area and the external ambient temperature during time period t, respectively. Let A be the internal temperature of the heating area at time t+1. a C represents the heating area of the park. a G represents the heat capacity per unit heating area. a This refers to the heat loss within a heating area per unit heating area due to temperature difference.
[0056] After establishing the above models, we can also establish the objective function for the two-sided coordinated control problem of the integrated energy system. For example, with the goal of minimizing the total cost of the integrated energy system, it can be expressed as:
[0057]
[0058] In the formula, It is the interaction cost between the integrated energy system and the external power grid in time period t. It is the gas purchase cost of the integrated energy system in time period t. This refers to the depreciation cost of charging and discharging energy storage devices. It is the load shedding cost of adjustable electrical load. This refers to the compensation cost for adjustable heat load, set up to ensure user thermal comfort. The various costs can be expressed using the following formulas:
[0059]
[0060]
[0061]
[0062]
[0063]
[0064] In the formula, Let P be the electricity price at time t for the park to purchase electricity from the external power grid and the electricity price for selling electricity to the external power grid. t grid Let t be the power exchange between the park and the external power grid. ρ is the unit price of natural gas purchased by the park's integrated energy system at time t. BES a1 is the depreciation cost coefficient for charging and discharging of the energy storage device, and a1 is the load shedding cost coefficient. The optimal heating zone temperature is set for users in the park. b1 represents the upper and lower limits of the allowable internal temperature of the heating area, respectively, and b1 is the compensation coefficient for the heat load of the park.
[0065] After constructing the objective function, pre-defined constraints that the integrated energy system's two-sided coordinated control problem needs to satisfy can also be constructed. These pre-defined constraints may include operating constraints for combined heat and power (CHP) units, operating constraints for electric energy storage devices, operating constraints for gas-fired boilers, operating constraints for electric boilers, power balance constraints, internal temperature constraints within the heating area, and interactive power constraints between the integrated energy system and the external power grid. Specifically, the operating constraints for CHP units are as follows:
[0066]
[0067] In the formula, These represent the minimum and maximum values of the output electrical power of the combined heat and power unit, respectively.
[0068] The specific operational constraints for energy storage devices are as follows:
[0069]
[0070]
[0071] In the formula, These are the minimum and maximum values of the charging power (or discharging power) of the energy storage device, respectively. These represent the lower and upper limits allowed for the State of Charge (SOC), respectively.
[0072] The specific operating constraints for gas-fired boilers are as follows:
[0073] In the formula, These are the minimum and maximum values of the output thermal power of the gas-fired boiler, respectively.
[0074] The specific power constraints between the integrated energy system and the external power grid are as follows:
[0075]
[0076] In the formula, These are the lower and upper limits of the power interaction between the integrated energy system and the external power grid, respectively.
[0077] Furthermore, the state space of the Markov decision process can be defined, i.e., real-time state data. Specifically, uncontrollable electrical load demand can be selected. Renewable energy power generation State of charge of energy storage devices Electricity purchase price Gas purchase price Temperature inside the heating area external ambient temperature And the time period t is used as a state-space variable, specifically:
[0078]
[0079] The actual electrical power consumed by the controllable electrical load, the electrical power output of the combined heat and power unit, the charging / discharging power of the electric energy storage device, the thermal power output of the electric boiler, and the thermal power output of the gas boiler are used as the action space variables, specifically:
[0080]
[0081] It should be noted that, since the actual electrical power consumed by the controllable electrical load, the electrical power output of the cogeneration unit, the operating power (charging power or discharging power) of the electric energy storage device, the thermal power output of the electric boiler, and the thermal power output of the gas boiler can be obtained, the thermal power output of the cogeneration unit and the interaction power between the system and the external power grid can be directly calculated. Therefore, the action space variables can be composed of the above five items.
[0082] In this embodiment of the invention, the optimization objective of the dual-sided coordinated control of the integrated energy system is to minimize the total cost. The objective of the Markov decision process is to maximize the long-term reward. Simultaneously, a penalty term is added to the immediate reward based on preset constraints to obtain the final decision reward function, which is expressed as follows:
[0083]
[0084] In the formula, rt As a reward value, This is a penalty item set up to reduce the excessive internal temperature of the heating area of the integrated energy system and the excessive power interaction with the external power grid. If the internal temperature of the heating area and the power interaction with the external power grid calculated after taking action are within the specified range, it is taken as 0; otherwise, a larger normal number is taken. 1 / 10000 is a corresponding scaling of the total cost.
[0085] In addition, a state-action value function, or simply state-action value function, can be defined for a Markov decision process. The state-action value function is as follows:
[0086]
[0087] In the formula, γ∈[0,1] is the discount factor, and E π (·) represents the expectation under policy π, which can be understood as state s and the corresponding action a, r. t+k Let Q be the reward value at time r+k. π (s, a) represents the state-action value under policy π, or simply the state-action value.
[0088] S2. Input the real-time status data into the pre-trained deep policy model. Through the policy network in the deep policy model, determine the decision action data of the integrated energy system. The decision action data includes the operating parameters of the equipment on the energy supply side and the controllable power consumption on the load side of the integrated energy system.
[0089] The deep policy model is trained based on an improved deep deterministic policy gradient algorithm, which combines a priority experience replay strategy and an L2 regularization strategy. The priority experience replay strategy is used to assign a corresponding priority to each sample in the experience replay pool, and the L2 regularization strategy is used to add penalty terms to the performance objective function and loss function to optimize the parameters of the deep policy model.
[0090] In one specific implementation, the training process of the deep policy model includes the following steps:
[0091] S21. Construct an initial model, which includes a value network, a policy network, a target value network, and a target policy network.
[0092] S22. Input the state data of each sample into the initial model, generate the predicted action data corresponding to the sample state data through the policy network, and obtain the predicted state action value corresponding to the predicted action data through the value network.
[0093] S23. Based on the predicted state action value and the corresponding target state action value, the loss function calculation result is obtained, and based on the predicted action data and the corresponding predicted state action value, the gradient function calculation result is obtained. The target state action value is calculated based on the target value network and the target policy network.
[0094] S24. Adjust the parameters of the value network according to the loss function calculation results, and adjust the parameters of the policy network according to the gradient function calculation results.
[0095] S25. Store the sample state data, the corresponding predicted action data, the predicted state action value, and the state data of the next moment into the experience replay pool, and determine the priority of the sample state data based on the absolute temporal difference error of the sample state data, so as to select sample state data from the experience replay pool through the priority to learn the value network and the policy network.
[0096] S26. Determine the deep policy model based on the trained policy network.
[0097] In step S21 above, an initial model is first constructed, including a value network, a policy network, a target value network, and a target policy network. The parameters of the value network and the policy network are initialized, and the parameters of the target value network and the target policy network are set to the same values as those of the value network. The target value network and the target policy network can be used to ensure the training of the value network and the policy network, avoiding excessive updates to the parameters of the value network and the policy network in each iteration.
[0098] Specifically, historical operating data of the integrated energy system can be collected in advance, and then sample status data can be constructed based on the historical operating data. The sample status data can include uncontrollable electrical load demand, renewable energy power generation, state of charge of energy storage devices, electricity purchase price, gas purchase price, internal temperature of heating area and external ambient temperature.
[0099] Furthermore, sample state data can be input into the initial model one by one. The input to the policy network in the initial model is an eight-dimensional state. The output is a five-dimensional motion. The input to the value network is the state s t and action a t The output is the state-action value (i.e., state-action value, which can be represented by Q(s)). t ,a t (This is represented by the symbol ). Therefore, the policy network has 8 neurons in the input layer and 5 neurons in the output layer, while the value network has 13 neurons in the input layer and 1 neuron in the output layer.
[0100] Specifically, the policy network can generate corresponding prediction action data based on the input sample state data (which may exclude time t, and time t can take a default value). The prediction action data may include the operating parameters of the equipment on the energy supply side of the integrated energy system and the controllable power consumption on the load side.
[0101] Regarding the above S22, optionally, the equipment operating parameters on the energy supply side in the predicted action data include the output electrical power of the combined heat and power unit, the output thermal power of the gas boiler, the output thermal power of the electric boiler, and the operating power of the electric energy storage device.
[0102] Each sample state data is input into the initial model, and the policy network generates the corresponding predicted action data through the policy network. This includes: inputting each sample state data into the initial model, and the policy network generating the corresponding predicted action data based on the sample state data, as well as the pre-built cogeneration unit model, gas boiler model, electric boiler model, electric energy storage device model, load-side heat load model, and preset constraints.
[0103] The output thermal power of the combined heat and power unit, as well as the interaction power between the system and the external power grid, can be calculated based on the output electrical power of the combined heat and power unit, the output thermal power of the gas boiler, the output thermal power of the electric boiler, and the operating power of the electric energy storage device.
[0104] In other words, the strategy network can input sample state data into models of combined heat and power units, gas boilers, electric boilers, electric energy storage devices, and load-side heat loads, and generate corresponding predicted action data under preset constraints. This method ensures that the predicted action data output by the strategy network meets the operational limitations of each device in the integrated energy system, guaranteeing the accuracy of the deep strategy model training process.
[0105] In one example, the load-side heat load model satisfies the following formula:
[0106]
[0107] In the formula, Let T be the heat load for time period t. t in T t out These represent the internal temperature of the heating area and the external ambient temperature at time t, respectively. Let A be the internal temperature of the heating area at time t+1. a C represents the heating area on the load side. a G represents the heat capacity per unit heating area. a Δt represents the heat loss within the heating area per unit heating area and per unit temperature difference, and Δt is the time interval.
[0108] It should be noted that the purpose of using the above-mentioned load-side heat load model to calculate the load-side heat load is as follows: the existing technology sets the heat load of the load side (such as buildings in a park) to a fixed value. However, in the embodiments of the present invention, considering that the human body is difficult to perceive temperature changes in a small range, and the user's demand for heating temperature is not fixed, in order to further achieve dual-side coordinated control with minimal total cost, the adjustable heat load can be calculated by combining the internal temperature of the heating area and the external ambient temperature, thereby achieving dual-side coordinated control, so as to further reduce costs while ensuring the heating comfort of users on the load side.
[0109] In one example, the preset constraints include operating constraints for combined heat and power units, operating constraints for electric energy storage devices, operating constraints for gas-fired boilers, operating constraints for electric boilers, power balance constraints, internal temperature constraints for the heating area, and interactive power constraints between the integrated energy system and the external power grid.
[0110] The specific formulas for each constraint can be found above. By constraining the combined heat and power units, electric energy storage devices, gas boilers, and electric boilers as described above, the rationality of the control of each device in the integrated energy system can be ensured. Furthermore, by constraining the internal temperature of the heating area, the heating comfort of users can be guaranteed.
[0111] After the policy network outputs the predicted action data, the sample state data and the corresponding predicted action data can be further input into the value network to obtain the predicted state-action value output by the value network. The predicted state-action value is the value calculated by the state-action value function.
[0112] Regarding the above S22, optionally, obtaining the predicted state action value corresponding to the predicted action data through the value network includes: inputting the sample state data and the corresponding predicted action data into the value network, so that the value network outputs the corresponding predicted state action value according to the state action value function;
[0113] Among them, the state action value function is constructed based on the decision reward function, which is constructed based on the objective function of the two-sided collaborative control. The objective function aims to minimize the total cost of the integrated energy system. The total cost includes the interaction cost between the integrated energy system and the external power grid, the gas purchase cost of the integrated energy system, the charging and discharging depreciation cost of the electric energy storage device, the load shedding cost of the adjustable electric load, and the compensation cost of the adjustable heat load.
[0114] The specific formulas for each cost can be found above. The state-action value function can be used to evaluate the quality of the coordinated control actions on both sides of the integrated energy transmission system. The larger the calculated result of the state-action value function, the better the predicted action data.
[0115] Specifically, the value network can evaluate the predicted action data output by the policy network to obtain predicted state action values. By constructing an objective function that minimizes the total cost of the integrated energy system, and then constructing a decision reward function based on the objective function, the state action value function is obtained. This achieves accurate construction of the state action value function, which facilitates the adjustment of the parameters of the value network and the policy network.
[0116] After the target value network outputs the predicted state-action value, the loss function can be calculated based on the predicted state-action value and the corresponding target state-action value. The gradient function can then be calculated based on the predicted action data and the corresponding predicted state-action value. The target state-action value can be obtained as follows: The state data for the next time step is obtained from the predicted action data. This next time step state data is input into the target policy network, causing the target policy network to output the action data for the next time step. Then, the next time step state data and the action data output by the target policy network are input into the target value network. The target state-action value is obtained based on the state-action value output by the target value network.
[0117] Specifically, the L2 regularization strategy in the improved deep deterministic policy gradient algorithm can be used to add a penalty term to the loss function, thereby optimizing parameter updates and improving the model's generalization ability in economic decision-making for integrated energy systems. Furthermore, the loss function is calculated by introducing the L2 regularization strategy. For example, after introducing L2 regularization into the loss function of the value network, the loss function and the corresponding gradient are expressed as follows:
[0118]
[0119]
[0120] In the formula, These are the calculated results of the loss function after adding the penalty term, and the corresponding gradient value, α. L Let be the regularization parameters for L2 regularization, and be the gradient function corresponding to the loss function, where:
[0121] L(θ Q )=E(y t -Q(s t a t |θ Q )) 2 ;
[0122] In the formula, L(θ) Q ) represents the original loss function (i.e., the loss function with L2 regularization), y t Let θ be the target state action value. QFor the parameters of the value network, Q(s) t a t |θ Q Let be the predicted state action value output by the value network, where the target state action value can be expressed as:
[0123] y t =r t +γQ′(s t+1 ,μ′(s t+1 |θ μ′ )|θ Q′ );
[0124] In the formula, S t+1 It is the next state that the integrated energy system transitions to after performing an action, θ Q′ θ represents the parameters of the target value network. μ′ For the parameters of the target policy network, μ′(s) t+1 |θ μ′ The target policy network is based on the input S. t+1 Output motion data, Q′(s) t+1 ,μ′(s t+1 |θ μ′ )|θ Q′ The target value network is based on the input S. t+1 And the state action values output by the action data output by the target policy network.
[0125] Regarding the above S23, optionally, the gradient function calculation result is obtained based on the predicted action data and the corresponding predicted state action value, including: substituting the predicted action data and the corresponding predicted state action value into the gradient function to obtain the gradient function calculation result.
[0126] The gradient function satisfies the following formula:
[0127]
[0128] In the formula, The gradient function calculation result is given, where N is the size of the mini-batch dataset, and s and a represent the state data and action data, respectively. i For the i-th sample state data, θ Q Let θ be the parameter of the value network. μ For the parameters of the policy network, The predicted action data output by the policy network. Let α be the predicted state action value output by the value network. L The regularization parameter for L2 regularization.
[0129] It should be noted that the above formula for calculating the gradient function is based on the improved gradient function after L2 regularization. The gradient function before the improvement could be:
[0130]
[0131] In the formula, J(θ) μ Let θ be the performance objective function of the policy network. μ Here are the parameters of the policy network. L2 regularization is introduced to improve the above performance objective function:
[0132]
[0133] Furthermore, the improved gradient function is expressed as:
[0134]
[0135] By introducing L2 regularization to improve the gradient algorithm for deep deterministic policies, the risk of overfitting during the training process of deep policy models can be reduced.
[0136] Furthermore, the parameters of the value network can be adjusted based on the loss calculation results, and the parameters of the policy network can be adjusted based on the gradient function calculation results. For example, the parameter update method of the value network is expressed as:
[0137]
[0138] In the formula, μ Q Let be the learning rate of the value network. The parameter update method of the policy network is expressed as:
[0139]
[0140] In the formula, μ μ Let be the learning rate of the policy network. Besides updating the parameters of the value network and the policy network, the parameters of the target value network and the target policy network can also be updated. The parameter update method for the target value network and the target policy network is expressed as follows:
[0141] θ Q′ ←τθ Q +(1-τ)θ Q′ ;
[0142] θ μ′ ←τθ μ +(1-τ)θ μ′ ;
[0143] In the formula, τ is the soft update coefficient, and τ << 1.
[0144] It should be noted that the policy network can also regenerate the predicted action data based on the gradient function calculation results, i.e., update the predicted action data. For example, the policy network can add random noise to the predicted action data to increase the network's exploration capability, as shown below:
[0145]
[0146] In the formula, For random noise, a t For the updated predicted action data, μ(s) t |θ μ ) represents the predicted action data initially generated by the policy network.
[0147] Specifically, after adjusting the parameters of the value network, the policy network, the target value network, and the target policy network, the next sample state data can be input into the initial model until all sample state data have been trained, i.e., repeating S22-S24.
[0148] In this embodiment of the invention, a priority experience replay strategy is introduced to improve the gradient algorithm for deep deterministic policies during model training. The trained sample state data can be stored in an experience replay pool for relearning.
[0149] Traditional experience replay mechanisms sample from the experience replay pool at the same frequency, without taking into account the different importance of each sample. In order to replay important samples more frequently, this invention introduces a priority experience replay strategy. By assigning different priorities to different samples, the relearning frequency of important samples is increased, thereby improving the decision accuracy of the model.
[0150] Specifically, the learned sample state data, along with the corresponding predicted action data and predicted state-action values, can be stored in the experience replay pool, and their priorities are determined based on the absolute temporal difference error. The priority reflects the importance of the samples in the experience replay pool, and the absolute temporal difference error is the absolute value of the difference between the target state-action value and the predicted state-action value, which can be represented by |δ... i | indicates, specifically:
[0151] |δ i |=|r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ )-Q(s i a i |θ Q )|;
[0152] For example, the sample state data in the experience replay pool can be sorted in descending order of absolute temporal difference error to obtain the sorting number of each sample state data. The reciprocal of the sorting number is then taken to obtain the priority of each sample state data. A larger absolute temporal difference error indicates higher quality of the corresponding prediction action data, lower importance of relearning, and consequently, lower priority.
[0153] Optionally, for the above S25, sample state data can be selected from the experience replay pool according to priority to learn the value network and policy network, including the following steps:
[0154] S251. Determine the sampling probability corresponding to each sample state data by using the priority corresponding to each sample state data and the number of priorities.
[0155] S252. Based on the sampling probability, the size of the empirical replay pool, and the correction amount corresponding to each sample state data, determine the importance sampling weight corresponding to each sample state data.
[0156] S253. Based on the importance sampling weights corresponding to the state data of each sample in the experience replay pool, the value network and policy network are learned.
[0157] Specifically, the sampling probability for each sample state data can be determined by the priority and number of priorities corresponding to each sample state data in the experience replay pool. For example, the sampling probability can be expressed as:
[0158]
[0159] In the formula, P i Let be the sampling probability of the i-th sample state data. The priority number, i.e., the number of priorities used, if An equal value of 0 indicates random sampling. Let i be the priority of the i-th sample state data under the priority number. This represents the sum of priorities for all sample state data under the specified priority level.
[0160] Furthermore, based on the sampling probability, the size of the empirical replay pool, and the correction amount, the importance sampling weight corresponding to each sample state data can be calculated, such as:
[0161]
[0162] In the formula, W i The importance sampling weights for the i-th sample state data are N. b P is the size of the experience replay pool, β is the correction amount used, and P is the value of the experience replay pool. i Let max be the sampling probability of the i-th sample state data.m W m The sampling weight is the one with the highest importance.
[0163] Furthermore, the value network and policy network can be learned using the importance sampling weights corresponding to the state data of each sample in the experience replay pool. The importance sampling weights can be used to calculate the original loss function for each state data sample in the experience replay pool. For example, the original loss function for each state data sample in the experience replay pool is as follows:
[0164]
[0165] Furthermore, the loss function calculation results corresponding to the state data of each sample in the experience replay pool can be obtained. Then, the parameters of the value network and the policy network can be adjusted using the loss function calculation results of the state data of each sample in the experience replay pool and the gradient function calculation results.
[0166] By employing the aforementioned priority experience replay strategy, greater weight can be assigned to important samples during training, thereby increasing the calculation results of the loss function. Compared to learning from randomly sampled samples, this approach can further improve the training efficiency of the model and the decision accuracy of the trained model.
[0167] In this embodiment of the invention, after training the deep policy model, the deep policy model can be used for decision-making. That is, real-time state data is input into the deep policy model to obtain the output results of the policy network in the deep policy model. The output results may include the actual power consumed by the controllable power consumption of the electrical load, the output power of the combined heat and power unit, the operating (charging or discharging) power of the electric energy storage device, the output thermal power of the electric boiler, and the output thermal power of the gas boiler. Then, the output thermal power of the combined heat and power unit and the interaction power between the system and the external power grid can be further calculated to obtain decision action data.
[0168] S3. Control the operation of each device on the power supply side based on the device's operating parameters, and control the power consumed on the load side of the integrated energy system based on the controllable power consumption.
[0169] Specifically, based on the equipment operating parameters in the decision-making action data, the operation of each device on the energy supply side of the integrated energy system can be controlled, and the power consumed by the integrated energy system on the load side can be controlled according to the controllable power consumption.
[0170] In this embodiment of the invention, a two-sided collaborative control action is obtained through the policy network in a pre-trained deep policy model. This method does not rely on the prediction or distribution information of renewable energy generation, electrical load, and heat load. Compared with the prior art, which predicts the future renewable energy generation and electrical load based on the current renewable energy generation and electrical load information and then makes decisions based on the predicted information, the method provided by this embodiment of the invention can solve the prediction error and avoid the problem of unreliable control scheme caused by the prediction error, thereby further ensuring the operating cost of the integrated energy system.
[0171] This invention offers the following technical advantages: By collecting real-time status data from an integrated energy system and inputting this data into a trained deep policy model, the policy network within the model generates decision-making data including equipment operating parameters on the energy supply side and controllable power consumption on the load side. This allows for control of the operation of each device on the energy supply side based on the equipment operating parameters, and control of the integrated energy system's energy supply to the load side based on controllable power consumption. This achieves coordinated optimization control of the load and energy supply sides of the integrated energy system. By dividing the load side electrical load into controllable and uncontrollable loads, and scheduling the controllable loads, the system can further reduce operating costs while ensuring responsiveness to load demand and meeting user thermal comfort. Furthermore, this method improves the deep deterministic policy gradient algorithm through a priority experience replay strategy and L2 regularization, thereby training a deep policy model that further enhances the accuracy of the model output and ensures the reliability of the dual-side coordinated control.
[0172] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 3 As shown, the electronic device 400 includes one or more processors 401 and memory 402.
[0173] The processor 401 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.
[0174] The memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may execute the program instructions to implement the integrated energy system dual-side cooperative control method of any embodiment of the present invention described above, and / or other desired functions. Various contents such as initial external parameters and thresholds may also be stored in the computer-readable storage medium.
[0175] In one example, the electronic device 400 may further include an input device 403 and an output device 404, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 403 may include, for example, a keyboard, a mouse, etc. The output device 404 may output various information to the outside, including warning messages, braking force, etc. The output device 404 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0176] Of course, for the sake of simplicity, Figure 3 Only some of the components of the electronic device 400 relevant to the present invention are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 400 may include any other suitable components depending on the specific application.
[0177] In addition to the methods and devices described above, embodiments of the present invention may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the integrated energy system dual-side cooperative control method provided in any embodiment of the present invention.
[0178] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0179] Furthermore, embodiments of the present invention may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the integrated energy system dual-side cooperative control method provided in any embodiment of the present invention.
[0180] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0181] It should be noted that the terminology used in this invention is for describing specific embodiments only and is not intended to limit the scope of this application. As shown in this specification, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may include the plural. The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.
[0182] It should also be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Unless otherwise expressly specified and limited, the terms "installed," "connected," "linked," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components. For those skilled in the art, the specific meaning of the above terms in the present invention can be understood according to the specific circumstances.
[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A dual-sided coordinated control method for an integrated energy system, characterized in that, include: S1. Collect real-time status data of the integrated energy system. The real-time status data includes uncontrollable electrical load demand, renewable energy power generation, state of charge of the energy storage device, electricity purchase price, gas purchase price, internal temperature of the heating area, and external ambient temperature. S2. Input the real-time status data into a pre-trained deep policy model, and determine the decision action data of the integrated energy system through the policy network in the deep policy model. The decision action data includes the equipment operating parameters on the energy supply side and the controllable power consumption on the load side of the integrated energy system. S3. Control the operation of each device on the power supply side based on the device operating parameters, and control the power consumed on the load side of the integrated energy system based on the controllable power consumption. The deep policy model is trained based on an improved deep deterministic policy gradient algorithm, which combines a priority experience replay strategy and an L2 regularization strategy. The priority experience replay strategy is used to assign a corresponding priority to each sample in the experience replay pool, and the L2 regularization strategy is used to add penalty terms to the performance objective function and loss function to optimize the parameters of the deep policy model. The training process of the deep policy model includes the following steps: S21. Construct an initial model, wherein the initial model includes a value network, a policy network, a target value network, and a target policy network; S22. Input the state data of each sample into the initial model, generate the predicted action data corresponding to the sample state data through the policy network, and obtain the predicted state action value corresponding to the predicted action data through the value network. S23. Based on the predicted state action value and the corresponding target state action value, the loss function calculation result is obtained, and based on the predicted action data and the corresponding predicted state action value, the gradient function calculation result is obtained. The target state action value is calculated based on the target value network and the target policy network. S24. Adjust the parameters of the value network according to the calculation result of the loss function, and adjust the parameters of the policy network according to the calculation result of the gradient function; S25. Store the sample state data, the corresponding predicted action data, the predicted state action value, and the state data at the next moment into the experience replay pool, and determine the priority of the sample state data based on the absolute temporal difference error of the sample state data, so as to select sample state data from the experience replay pool according to the priority to learn the value network and the policy network. S26. Determine the deep policy model based on the trained policy network.
2. The method according to claim 1, characterized in that, The predicted action data includes the equipment operating parameters on the energy supply side and the controllable power consumption on the load side. The equipment operating parameters on the energy supply side in the predicted action data include the output power of the combined heat and power unit, the output heat power of the gas boiler, the output heat power of the electric boiler, and the operating power of the electric energy storage device. The step of inputting the state data of each sample into the initial model and generating the predicted action data corresponding to the sample state data through the policy network includes: Each sample state data is input into the initial model. The strategy network generates corresponding prediction action data based on the sample state data, as well as pre-built models of cogeneration units, gas boilers, electric boilers, electric energy storage devices, load-side heat loads, and preset constraints.
3. The method according to claim 2, characterized in that, The load-side heat load model satisfies the following formula: ; In the formula, Let be the heat load at time t. , They are respectively time t The internal temperature of the heating area and the external ambient temperature, The temperature inside the heating area at time t+1. This refers to the heating area on the load side. This refers to the heat capacity per unit heating area. This refers to the heat loss within a heating area per unit heating area and per unit temperature difference. For time intervals.
4. The method according to claim 2, characterized in that, The preset constraints include operating constraints for combined heat and power units, operating constraints for electric energy storage devices, operating constraints for gas-fired boilers, operating constraints for electric boilers, power balance constraints, internal temperature constraints for heating areas, and interactive power constraints between the integrated energy system and the external power grid.
5. The method according to claim 1, characterized in that, The step of obtaining the predicted state action value corresponding to the predicted action data through the value network includes: The sample state data and the corresponding predicted action data are input into the value network, so that the value network outputs the corresponding predicted state action value according to the state action value function. The state action value function is constructed based on the decision reward function, which is constructed based on the objective function of the two-sided collaborative control, and the objective function aims to minimize the total cost of the integrated energy system. The total cost includes the interaction cost between the integrated energy system and the external power grid, the gas purchase cost of the integrated energy system, the depreciation cost of charging and discharging the electric energy storage device, the load shedding cost of the adjustable electrical load, and the compensation cost of the adjustable heat load.
6. The method according to claim 1, characterized in that, The step of obtaining the gradient function calculation result based on the predicted action data and the corresponding predicted state action value includes: Substitute the predicted action data and the corresponding predicted state action value into the gradient function to obtain the gradient function calculation result; The gradient function satisfies the following formula: ; In the formula, The gradient function is the result of the calculation, and N is the size of the mini-batch dataset. , These represent state data and action data, respectively. For the i-th sample state data, The parameters of the value network, The parameters of the policy network are... The predicted action data output by the policy network. The predicted state action value output by the value network. The regularization parameter for L2 regularization.
7. The method according to claim 1, characterized in that, The step of selecting sample state data from the experience replay pool by priority to learn the value network and the policy network includes the following steps: S251. Determine the sampling probability corresponding to each sample state data by using the priority corresponding to each sample state data and the number of priorities. S252. Based on the sampling probability corresponding to each sample state data, the size of the empirical replay pool, and the correction amount, determine the importance sampling weight corresponding to each sample state data. S253. Based on the importance sampling weights corresponding to the state data of each sample in the experience replay pool, the value network and the policy network are learned.
8. An electronic device, characterized in that, The electronic device includes: Processor and memory; The processor executes the steps of the integrated energy system dual-side cooperative control method as described in any one of claims 1 to 7 by calling the program or instructions stored in the memory.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform the steps of the integrated energy system dual-side coordinated control method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
TD3-based comprehensive energy system source-load collaborative operation optimization method
CN114462696A
Distributed energy optimization scheduling method and device based on mixed learning
CN116451880A