A collaborative control method for multiple intelligent agents in an integrated energy system and the energy system
By introducing a collaborative control method involving multiple intelligent agents into the integrated energy system, and using neural network models to predict and control electrical parameters, the problem of supply and demand mismatch was solved, a stable energy supply balance was achieved, and the phenomenon of wind and solar curtailment was reduced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2026-03-13
Smart Images

Figure CN116384250B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, specifically to a collaborative control method and energy system for multiple intelligent agents in an integrated energy system. Background Technology
[0002] Currently, my country is in the exploratory stage of clean energy structure transformation. To reduce carbon emissions, the integrated energy system introduces a large amount of renewable energy on the "source side" to make energy input cleaner, while the "load side" achieves clean energy consumption through technologies such as electricity substitution. However, the randomness and volatility of renewable energy and human energy consumption behavior inevitably introduce uncertainty into the system, leading to a supply-demand mismatch and causing the phenomenon of "wind and solar curtailment".
[0003] In related technologies, a common approach to achieve a cleaner energy structure transformation is to add energy storage devices to the energy system to balance the "source side" and "load side" of the integrated energy system. However, due to economic and safety constraints, the capacity of energy storage devices is usually limited, making it impossible to guarantee the energy supply and demand balance between the "source side" and "load side" of the integrated energy system. Summary of the Invention
[0004] Therefore, the technical problem to be solved by the present invention is to overcome the technical defects in the prior art that cannot guarantee the energy supply and demand balance of the integrated energy system, thereby providing a collaborative control method and energy system for multiple intelligent agents in an integrated energy system.
[0005] In a first aspect, embodiments of the present invention provide a collaborative control method for multiple intelligent agents in an integrated energy system. The integrated energy system includes a first predictive intelligent agent, a second predictive intelligent agent, and a first control intelligent agent. The first control intelligent agent includes at least one adjustable unit. The method includes: the first control intelligent agent receiving a first electrical energy parameter from the first predictive intelligent agent and a second electrical energy parameter from the second predictive intelligent agent, respectively. The first electrical energy parameter is an output electrical parameter predicted by a first neural network model, and the second electrical energy parameter is a consumption electrical parameter predicted by a second neural network model; the first control intelligent agent processes the first electrical energy parameter and the second electrical energy parameter through a third neural network model and outputs at least one control parameter; the first control intelligent agent controls the operation of at least one adjustable unit through the at least one control parameter.
[0006] In conjunction with the first aspect, in one possible implementation of the first aspect, the third neural network model includes: a policy network and a target network. Before the first control agent processes the first and second electrical energy parameters through the third neural network model, the model further includes: acquiring the third neural network model. Acquiring the third neural network model includes: establishing datasets for the policy network and the target network, the datasets including: training data and rewards corresponding to the control agent; processing the training dataset through the policy network to obtain a first fitted value of the policy network; processing the training dataset through the target network to obtain a second fitted value of the target network; solving for the error between the first and second fitted values using a mean squared error loss function; and determining the policy parameters of the output policy network using the error and a preset threshold.
[0007] In conjunction with the first aspect, in one possible implementation of the first aspect, establishing a dataset for a policy network and a target network includes: obtaining a first state of the controlling agent; processing the first state through the policy network to obtain the action value of the controlling agent; establishing a simulation model corresponding to the adjustable unit; inputting the action value into the simulation model to obtain a second state of the controlling agent; calculating the reward corresponding to the controlling agent based on the second state; and using the first state, action value, and second state as training data, wherein the dataset includes training data and reward.
[0008] In conjunction with the first aspect, in one possible implementation of the first aspect, the reward corresponding to the control agent is calculated based on the second state, including: determining the type of reward; and calculating the reward corresponding to the control agent based on the second state and the type of reward.
[0009] In conjunction with the first aspect, in one possible implementation of the first aspect, calculating the reward corresponding to the control agent includes: determining the energy cost reward corresponding to the first control agent through a first relation and a second relation.
[0010] The first relation is:
[0011]
[0012] The second relation is:
[0013]
[0014] Where, r cost Let M represent the energy cost reward, i represent the number of first-level control agents, and j represent the number of first-level control agents that can engage in electricity sales, j∈i. n W represents the electricity purchase / sale cost of each first-controlling intelligent agent. n W represents the electrical energy corresponding to each first controlling agent. n A value greater than zero indicates an electricity purchase.n A value less than zero indicates electricity sales activity, P pur P represents the current electricity price. sell This indicates the current electricity price.
[0015] In conjunction with the first aspect, in one possible implementation of the first aspect, the integrated energy system further includes: a second control agent, which includes at least one adjustable unit, calculates a reward corresponding to the control agent, and further includes: determining a flexible load control reward corresponding to the second control agent through a third relation and a fourth relation.
[0016] The third relation is:
[0017]
[0018] The fourth relation is:
[0019]
[0020] Where, r EV This represents the reward for flexible load control, where t represents the current time. set Indicates the preset time. W represents the state of charge of the load currently connected to the second control agent. EV This indicates the charge / discharge level of the second controlling agent. This represents the total capacity of the load connected to the second control agent. This indicates the state of charge of the load connected to the second control agent at the previous moment.
[0021] In conjunction with the first aspect, in one possible implementation of the first aspect, the integrated energy system further includes: a third control agent, which includes at least one adjustable unit, calculates a reward corresponding to the control agent, and further includes: determining a comfort reward corresponding to the third control agent through a fifth relation.
[0022] The fifth relation is:
[0023]
[0024] Where, r com Indicates comfort reward, T max T represents the maximum temperature value within the comfortable temperature range. min This represents the minimum temperature value within the comfortable temperature range. represents the set temperature value of the k-th third control agent, and l represents the number of third control agents.
[0025] Secondly, embodiments of the present invention provide a collaborative control device for multiple intelligent agents in an integrated energy system. The integrated energy system includes a first predictive intelligent agent, a second predictive intelligent agent, and a first control intelligent agent. The first control intelligent agent includes at least one adjustable unit. The device includes: a receiving unit for the first control intelligent agent to receive a first electrical energy parameter from the first predictive intelligent agent and a second electrical energy parameter from the second predictive intelligent agent, wherein the first electrical energy parameter is an output electrical parameter predicted by a first neural network model, and the second electrical energy parameter is a consumption electrical parameter predicted by a second neural network model; an output unit for the first control intelligent agent to process the first electrical energy parameter and the second electrical energy parameter through a third neural network model and output at least one control parameter; and a control unit for the first control intelligent agent to control the operation of at least one adjustable unit through at least one control parameter.
[0026] Thirdly, embodiments of the present invention provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor;
[0027] The memory stores instructions that can be executed by at least one processor. When the instructions are executed by at least one processor, the electronic device performs a cooperative control method for multiple intelligent agents in an integrated energy system as described in any embodiment of the first aspect.
[0028] Fourthly, embodiments of the present invention provide an integrated energy system, comprising a first predictive agent, a second predictive agent, and a first control agent. The first control agent includes at least one adjustable unit. The first predictive agent collects first meteorological data and inputs it into a first neural network model to obtain a first electrical energy parameter, which is a predicted output electrical parameter. The second predictive agent collects second meteorological data and inputs it into a second neural network model to obtain a second electrical energy parameter, which is a predicted consumption electrical parameter. The first control agent receives the first electrical energy parameter from the first predictive agent and the second electrical energy parameter from the second predictive agent. The first control agent processes the first and second electrical energy parameters through a third neural network model and outputs at least one control parameter. The first control agent controls the operation of at least one adjustable unit using the at least one control parameter.
[0029] This invention provides a collaborative control method and energy system for multiple intelligent agents in an integrated energy system. The integrated energy system includes a first predictive intelligent agent, a second predictive intelligent agent, and a first control intelligent agent. The first control intelligent agent includes at least one adjustable unit. The method includes: the first control intelligent agent receiving a first electrical energy parameter from the first predictive intelligent agent and a second electrical energy parameter from the second predictive intelligent agent, respectively. The first electrical energy parameter is a predicted output electrical energy parameter, and the second electrical energy parameter is a predicted consumption electrical energy parameter. The first control intelligent agent processes the first and second electrical energy parameters through a third neural network model and outputs at least one control parameter. The first control intelligent agent controls the operation of at least one adjustable unit through the at least one control parameter. In this process, by introducing a control intelligent agent, a first predictive intelligent agent, and a second predictive intelligent agent into the integrated energy system, the control intelligent agent, the first predictive intelligent agent, and the second predictive intelligent agent respectively represent the energy storage side, the source side, and the load side in the integrated energy system. Based on the predicted output electrical energy and the predicted consumption electrical energy, the third neural network model determines at least one control parameter for the control intelligent agent, thereby achieving energy balance in the integrated energy system through joint regulation of the energy storage side, the source side, and the load side. Attached Figure Description
[0030] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0031] Figure 1 This is an example diagram illustrating an application scenario of an integrated energy system provided by an embodiment of the present invention;
[0032] Figure 2 This is an example diagram illustrating another application scenario of an integrated energy system provided by an embodiment of the present invention;
[0033] Figure 3 A flowchart illustrating a specific example of a collaborative control method for multiple intelligent agents in an integrated energy system provided by an embodiment of the present invention;
[0034] Figure 4 A flowchart illustrating another specific example of a collaborative control method for multiple intelligent agents in an integrated energy system provided by an embodiment of the present invention;
[0035] Figure 5 A schematic diagram illustrating a specific example of a collaborative control device for multiple intelligent agents in an integrated energy system, provided in an embodiment of the present invention.
[0036] Figure 6 This is a structural example diagram of an electronic device in an embodiment of the present invention. Detailed Implementation
[0037] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] In the description of this invention, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0039] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can also refer to the internal connection of two components; and they can refer to a wireless connection or a wired connection. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0040] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0041] An integrated energy system provided in this embodiment of the invention, such as Figure 1 As shown, the integrated energy system includes a first predictive agent, a second predictive agent, and at least one control agent, with each control agent including at least one adjustable unit.
[0042] Specifically, the first predictive agent includes: first predictive agent 11 and first predictive agent 12. The second predictive agent includes: second predictive agent 211 and second predictive agent 221. The control agent includes: control agent 212, control agent 213, control agent 221, control agent 223, control agent 31, control agent 32, and control agent 33.
[0043] Specifically, such as Figure 2As shown, the first predictive agent 11 and the first predictive agent 12 represent a wind power generation device and a solar power generation device, respectively, for generating electrical energy; the second predictive agent 211 and the second predictive agent 221 represent an office building and a residential building, respectively, utilizing the electrical energy generated by the first predictive agent 11 and / or the first predictive agent 12; the control agent 212 and the control agent 221 represent a heat pump, respectively utilizing the electrical energy generated by the first predictive agent 11 and / or the first predictive agent 12 to provide heat energy to the office building or the residential building; the control agent 211... 13 and control agent 223 represent charging piles, respectively using the electrical energy generated by the first predictive agent 11 and / or the first predictive agent 12 to charge the equipment in the office building area 21 or the residential building area 22, or the equipment in the office building area 21 or the residential building area 22 to charge the charging piles; control agents 31, 32 and 33 represent various energy storage devices, including: supercapacitors, lithium batteries, compressed air energy storage devices, etc., used to store the electrical energy generated by the first predictive agent 11 and / or the first predictive agent 12.
[0044] In one alternative embodiment, the control agent includes: a first control agent, and the integrated energy system includes a first predictive agent, a second predictive agent, and at least one first control agent, each control agent including at least one adjustable unit.
[0045] Specifically, the first control agent is one or more of control agent 31, control agent 32, and control agent 33.
[0046] In this approach, the integrated energy system comprises a first predictive agent, a second predictive agent, and a first control agent. The first predictive agent acts as the source side, generating electrical energy; the second predictive agent acts as the load side, consuming electrical energy; and the first control agent acts as the energy storage side, storing electrical energy. Thus, by coordinating the regulation of the energy storage side, source side, and load side within the integrated energy system, the energy supply balance of the integrated energy system is achieved.
[0047] In one optional embodiment, the control agent includes: a first control agent and a second control agent. The integrated energy system includes a first predictive agent, a second predictive agent, at least one first control agent and at least one second control agent. Each first control agent includes at least one adjustable unit.
[0048] Specifically, the second control agent is one or more of control agent 213 and control agent 223.
[0049] In this approach, the integrated energy system comprises a first predictive agent, a second predictive agent, a first control agent, and a second control agent. The first predictive agent acts as the source side, generating electrical energy; the second predictive agent and the second control agent act as the load side, consuming electrical energy; and the first control agent acts as the energy storage side, storing electrical energy. Thus, by jointly regulating the energy storage side, source side, and load side of the integrated energy system, energy supply balance is achieved.
[0050] In one alternative embodiment, the control agent includes: a first control agent, a second control agent, and a third control agent; the integrated energy system includes a first predictive agent, a second predictive agent, at least one first control agent, at least one second control agent, and at least one third control agent, each control agent including at least one adjustable unit.
[0051] Specifically, the third control agent is one or more of control agent 212 and control agent 221.
[0052] In this approach, the integrated energy system comprises a first predictive agent, a second predictive agent, a first control agent, a second control agent, and a third control agent. The first predictive agent acts as the source side, generating electrical energy; the second, second, and third control agents act as the load side, consuming electrical energy; and the first control agent acts as the energy storage side, storing electrical energy. Thus, by coordinating the regulation of the energy storage side, source side, and load side within the integrated energy system, a balance in energy supply is achieved.
[0053] The first predictive agent, the second predictive agent, and the control agent can transmit and exchange information through the integrated energy hub 4 or other means, and the present invention does not impose specific limitations on this. For ease of explanation, the integrated energy hub 4 is used in the figure. It should be understood that the number of the first predictive agent, the second predictive agent, and the control agent, the connection method, the method of information transmission and exchange, and the physical devices corresponding to the first predictive agent, the second predictive agent, and the control agent include, but are not limited to, those shown in the figure.
[0054] The first predictive agent is used to collect first meteorological data, input the first meteorological data into the first neural network model, and obtain the first electrical energy parameter, which is the predicted output electrical parameter.
[0055] The second predictive agent is used to collect second meteorological data, input the second meteorological data into the second neural network model, and obtain the second electrical energy parameter, which is the predicted electrical consumption parameter.
[0056] Specifically, the first meteorological data and the second meteorological data represent meteorological data related to the power generation method. When the first predictive agent uses wind power generation, the first meteorological data can be data such as wind speed, wind direction, and outdoor temperature. When the second predictive agent uses solar power generation, the second meteorological data can be data such as solar radiation intensity, altitude angle, and outdoor temperature.
[0057] Specifically, before inputting the first meteorological data into the first neural network model, the method further includes: acquiring the first neural network model; and before inputting the second meteorological data into the second neural network model, the method further includes: acquiring the second neural network model.
[0058] Specifically, the first neural network model and the second neural network model can be the same neural network model or different neural network models; this invention does not impose specific limitations on this. For ease of explanation, taking the example where both the first and second neural network models use the BP neural network algorithm, the input features of the BP neural network algorithm include: wind speed, wind direction, outdoor temperature, and historical power generation power included in the first predictive agent 11 as input features, with wind turbine power generation power as the label value; solar radiation intensity, altitude angle, outdoor temperature, and historical power generation power included in the first predictive agent 12 as input features, with photovoltaic power generation power as the label value; date, time, number of people, outdoor temperature, wind speed, door and window opening status, and indoor temperature included in the second predictive agent 211 as input features, with office building energy consumption as the label value; and date, time, number of people, outdoor temperature, wind speed, door and window opening status, and indoor temperature included in the second predictive agent 221 as input features, with residential energy consumption as the label value.
[0059] Specifically, to obtain the BP neural network algorithm, such as Figure 3 As shown, it includes:
[0060] S11. Initialize connection weights and thresholds.
[0061] S12, Input normalized training data.
[0062] S13. Calculate the input and output of each neuron in the intermediate layer.
[0063] S14. Calculate the input and output of each neuron in the output layer.
[0064] S15. Calculate the error between the output layer result and the label value.
[0065] S16. Calculate the error of each neuron in the intermediate layer.
[0066] S17. Update the connection weights and thresholds between each layer.
[0067] S18. Output the result when the error is less than the set threshold.
[0068] Specifically, the error refers to the average absolute percentage error between the output result and the label value. Training ends when the average absolute percentage error between the agent's output result and the label value is less than 10%, and the prediction result is output. It should be understood that using the BP neural network algorithm for prediction and determining the prediction result is a relatively mature technology, and this invention will not elaborate on it further. Specifically, when the error exceeds the set threshold, the process returns to step S12 and repeats steps S12 to S18.
[0069] The control agent receives a first electrical energy parameter from the first predictive agent and a second electrical energy parameter from the second predictive agent.
[0070] The control agent processes the first and second electrical energy parameters through a third neural network model and outputs at least one control parameter.
[0071] The control agent controls the operation of at least one adjustable unit through at least one control parameter.
[0072] By implementing this embodiment, a control agent, a first predictive agent, and a second predictive agent are introduced into the integrated energy system. The control agent, the first predictive agent, and the second predictive agent respectively represent the energy storage side, the source side, and the load side in the integrated energy system. The generated electrical energy and the consumed electrical energy are predicted by a first neural network model. At least one control parameter of the control agent is determined by a third neural network model. In this process, the first predictive agent determines the electrical energy generated by the source side, the second predictive agent determines all or part of the electrical energy consumed by the load side, and the control agent determines the electrical energy stored by the energy storage side, or the electrical energy stored by the energy storage side and part of the electrical energy consumed by the load side. Thus, the energy supply balance of the integrated energy system is achieved by jointly regulating the energy storage side, the source side, and the load side.
[0073] This embodiment provides a collaborative control method for multiple intelligent agents in an integrated energy system, such as... Figure 4 As shown, the integrated energy system includes a first predictive agent, a second predictive agent, and a first control agent. The first control agent includes at least one adjustable unit. The control method includes the following steps:
[0074] S21: The first control agent receives a first electrical energy parameter from the first predictive agent and a second electrical energy parameter from the second predictive agent. The first electrical energy parameter is the output electrical parameter predicted by the first neural network model, and the second electrical energy parameter is the consumption electrical parameter predicted by the second neural network model.
[0075] Specifically, the first electrical energy parameter sent by the first predictive agent is obtained by collecting first meteorological data and inputting the first meteorological data into the first neural network model. For details, please refer to the relevant description of the first neural network model in the above embodiments, which will not be repeated here.
[0076] Specifically, the second electrical energy parameter sent by the second predictive agent is obtained by collecting second meteorological data and inputting the second meteorological data into the second neural network model. For details, please refer to the relevant description of the second neural network model in the above embodiments, which will not be repeated here.
[0077] S22: The first control agent processes the first and second electrical energy parameters through the third neural network model and outputs at least one control parameter.
[0078] Specifically, the third neural network model includes a policy network and a target network. The policy network is used to fit the input parameters and determine the fitting result corresponding to the input parameters. The target network is used to judge whether the fitting result determined by the policy network is correct.
[0079] S23: The first control agent controls the operation of at least one adjustable unit through at least one control parameter.
[0080] Specifically, the adjustable unit included in the first control agent, when controlled by control parameters determined by the third neural network model, is used to adjust the energy storage capacity of one or more of the supercapacitors, lithium batteries, and compressed air energy storage devices corresponding to the first control agent.
[0081] In one optional embodiment, the control agent further includes a second control agent and a third control agent. The adjustable unit included in the second control agent, when controlled by control parameters determined by the third neural network model, is used to adjust the charging or discharging amount of the charging pile corresponding to the second control agent, i.e., to adjust the amount of electricity purchased or sold by the second control agent. The adjustable unit included in the third control agent, when controlled by control parameters determined by the third neural network model, is used to adjust the power consumption of the heat pump corresponding to the third control agent, i.e., to adjust the amount of heat energy provided by the heat pump to the second predictive agent.
[0082] By implementing this embodiment, a control agent, a first predictive agent, and a second predictive agent are introduced into the integrated energy system. The control agent, the first predictive agent, and the second predictive agent respectively represent the energy storage side, the source side, and the load side in the integrated energy system. Based on the predicted electricity output and the predicted electricity consumption, at least one control parameter of the control agent is determined through a third neural network model. In this way, the energy supply balance of the integrated energy system is achieved by jointly regulating the energy storage side, the source side, and the load side.
[0083] In one optional implementation, to ensure that the control parameters output by the third neural network model conform to the actual operating conditions, the third neural network model needs to be pre-trained. That is, before the control agent processes the first and second electrical energy parameters through the third neural network model, the method further includes: acquiring the third neural network model.
[0084] Obtain the third neural network model, including:
[0085] (1) Establish the datasets for the policy network and the target network. The datasets include: training data and rewards corresponding to the control agent.
[0086] Specifically, the third neural network may use the Deep Deterministic Policy Gradient (DDPG) algorithm or other neural network algorithms, as long as the third neural network model includes a policy network and a target network, and the target network is used to determine whether the output of the policy network is correct.
[0087] In one alternative implementation, a dataset for both the policy network and the target network is established, including:
[0088] Obtain the first state of the controlling agent.
[0089] Specifically, the first state of the controlling agent refers to the initial state of the controlling agent, in order to... Figure 1 , Figure 2 Taking the application scenario shown as an example, the first state of each control agent is denoted as: When the controlling agent is the third controlling agent, the corresponding state of the third controlling agent includes: date, time, wind speed, wind direction, outdoor temperature, indoor temperature, solar radiation intensity, predicted wind turbine power generation, predicted photovoltaic power generation, current heat pump setting, predicted office building energy consumption, predicted residential energy consumption, SOC of each energy storage device, and electricity price.
[0090] When the controlling agent is the second controlling agent, the corresponding state of the second controlling agent includes: date, time, wind speed, wind direction, outdoor temperature, indoor temperature, solar radiation intensity, predicted wind turbine power generation, predicted photovoltaic power generation, current heat pump setting, predicted office building energy consumption, predicted residential energy consumption, SOC of each energy storage device, electricity price, and heat pump temperature setting value.
[0091] When the controlling agent is the first controlling agent, the corresponding state of the first controlling agent includes: date, time, wind speed, wind direction, outdoor temperature, indoor temperature, solar radiation intensity, predicted wind turbine power generation, predicted photovoltaic power generation, current heat pump setting, predicted office building energy consumption, predicted residential energy consumption, SOC of each energy storage device, electricity price, heat pump temperature setting value, and charging pile working status.
[0092] As can be seen, in this embodiment, when the controlling agent includes other controlling agents besides the first controlling agent, the state of the first controlling agent includes the states of the second and third controlling agents, and the state of the second controlling agent includes the state of the first controlling agent. Essentially, it is necessary to first determine the state of the third controlling agent, then determine the state of the second controlling agent based on the state of the third controlling agent, and finally determine the state of the first controlling agent based on the states of the third and second controlling agents. That is, the third controlling agent, as a physical device that engages in electricity purchase, can only consume electrical energy; the second controlling agent, as a physical device capable of both purchasing and selling electricity, can consume and provide electrical energy; and the first controlling agent, as an energy storage device, can store electrical energy.
[0093] The first state is processed by the policy network to obtain the action value of the controlling agent.
[0094] Specifically, processing the first state through the policy network means inputting all corresponding states into the current policy network to calculate the value of the action. Figure 1 , Figure 2 Taking the application scenario shown as an example, the action values of each control agent are recorded as follows:
[0095] Specifically, when the controlling agent is the third controlling agent, its action value refers to the heat pump temperature setpoint. When the controlling agent is the second controlling agent, its action value refers to the charging pile's electricity value, i.e., the electricity purchased or sold. When the controlling agent is the first controlling agent, its action value refers to the charging and discharging amount of the energy storage device.
[0096] Establish a simulation model corresponding to the adjustable unit.
[0097] Specifically, establishing a simulation model corresponding to at least one adjustable unit in the control agent is a relatively mature technology, and this invention will not elaborate on it further. As long as the established simulation model can reflect the actual working conditions of the first control agent, the second control agent, and the third control agent, it is sufficient.
[0098] Input the action values into the simulation model to obtain the second state of the control agent.
[0099] Specifically, the second state of the controlling agent refers to the state determined by each controlling agent at the next moment after performing the aforementioned actions, in order to... Figure 1 , Figure 2 Taking the application scenario shown as an example, the second state of each control agent is denoted as...
[0100] Based on the second state, calculate and control the corresponding reward for the agent.
[0101] Specifically, the reward corresponding to the controlling agent is represented by the following formula:
[0102] r = r cost ×ω cost +r EV ×ω EV +r com ×ω com
[0103] Where r represents the reward at the current moment, r cost Indicates energy cost incentive, r EV Indicates a reward for flexible load control, r com Indicates comfort reward, ω cost ω represents the weight of the energy cost incentive. EV ω represents the weight of the flexible load control reward. com This represents the weight of the comfort reward. Specifically, ω cost ω EV ω com The sum of all values equals 1, and the value of each weight ranges from 0 to 1, set according to specific operating conditions; this invention does not impose specific limitations on this. When the integrated energy system does not include a second control agent, the weight of the flexible load control reward is 0. When the integrated energy system does not include a third control agent, the weight of the comfort reward is 0.
[0104] In one alternative implementation, based on the second state, the reward corresponding to the controlling agent is calculated, including:
[0105] Determine the type of reward.
[0106] Specifically, the types of rewards include one or more of the following: energy cost rewards, flexible load control rewards, and comfort rewards, with the type of reward determined based on the control agent. Each control agent needs to calculate the energy cost reward; the flexible load control reward corresponds to the second control agent, meaning that if the control agent includes the second control agent, then the flexible load control reward needs to be calculated; the comfort reward corresponds to the third control agent, meaning that if the control agent includes the third control agent, then the comfort reward needs to be calculated.
[0107] Based on the second state and the type of reward, calculate and control the corresponding reward for the agent.
[0108] In one alternative implementation, calculating the reward corresponding to the control agent includes:
[0109] The energy cost reward corresponding to the first control agent is determined by the first and second relational expressions.
[0110] The first relation is:
[0111]
[0112] The second relation is:
[0113]
[0114] Where, r cost Let M represent the energy cost reward, i represent the number of first-level control agents, and j represent the number of first-level control agents that can engage in electricity sales, j∈i. n W represents the electricity purchase / sale cost of each first-controlling intelligent agent. n W represents the electrical energy consumed by each first control agent. n A value greater than zero indicates an electricity purchase. n A value less than zero indicates electricity sales activity, P pur P represents the current electricity price. sell This indicates the current electricity price.
[0115] Specifically, with Figure 1 , Figure 2 Taking the illustrated application scenario as an example, there are 7 control agents, including: first control agent 31, first control agent 32, first control agent 33, second control agent 213, second control agent 223, third control agent 212, and third control agent 222. Since each control agent needs to calculate energy costs, the first and second relational expressions are written as follows:
[0116] First relation:
[0117]
[0118] The second relation is:
[0119]
[0120] Specifically, the value of i in the above formula is the number of control agents. Since the number of first control agents is 2, and the first control agent only has the behavior of purchasing electricity, that is, it can only consume electricity, the value of j in the above formula is 2, that is, the sum of the number of second and third control agents after excluding the first control agent.
[0121] In one alternative embodiment, the integrated energy system further includes: a second control agent, which includes at least one adjustable unit.
[0122] The rewards corresponding to computational and control agents also include:
[0123] The flexible load control reward corresponding to the second control agent is determined by the third and fourth relations.
[0124] The third relation is:
[0125]
[0126] The fourth relation is:
[0127]
[0128] Where, r EV This represents the reward for flexible load control, where t represents the current time. set Indicates the preset time. W represents the state of charge of the load currently connected to the second control agent. EV This indicates the charge / discharge level of the second controlling agent. This represents the total capacity of the load connected to the second control agent. This indicates the state of charge of the load connected to the second control agent at the previous moment.
[0129] Specifically, with Figure 1 , Figure 2 Taking the application scenario shown as an example, t set Used to indicate the time when office workers leave work, the time when residents of residential buildings are not in their residential areas, or other corresponding times when the first predictive agent does not consume power. Used to indicate the current state of charge of the electric vehicle. Used to indicate the total battery capacity of an electric vehicle.
[0130] In one alternative embodiment, the integrated energy system further includes: a third control agent, which includes at least one adjustable unit.
[0131] The rewards corresponding to computational and control agents also include:
[0132] The comfort reward corresponding to the third control agent is determined through the fifth relation.
[0133] The fifth relation is:
[0134]
[0135] Where, r com Indicates comfort reward, T max T represents the maximum temperature value within the comfortable temperature range. min This represents the minimum temperature value within the comfortable temperature range. represents the set temperature value of the k-th third control agent, and l represents the number of third control agents.
[0136] Specifically, with Figure 1 , Figure 2 Taking the application scenario shown as an example, if there are two third control agents, namely third control agent 212 and third control agent 222, then the above formula can be written as:
[0137]
[0138] Specifically, in the above formula, Used to represent the set temperature of the third control agent 212 Used to indicate the set temperature of the third control agent 222.
[0139] The first state, action value, and second state are used as training data. The dataset includes training data and rewards.
[0140] Specifically, with Figure 1 , Figure 2 Taking the application scenario shown as an example, using the first state, action value, and second state as training data means that... As a training data point, the second state at the current moment is used as the first state at the next moment. Thus, each training data point, along with its corresponding reward—that is, the reward at the current moment—is aggregated with the training data to form a dataset. The training process of the third neural network is equivalent to iteratively finding the maximum value of the reward function. The closer the reward function is to 1, the better the control performance of the integrated energy system meets the requirements.
[0141] (2) Process the training dataset through the policy network to obtain the first fitted value of the policy network.
[0142] Specifically, a defined training dataset is input into the policy network to obtain the fitted value of the policy network. For example, the training dataset is input into the current policy network of the DDPG algorithm to determine the Q-value of each agent, and the sum of the Q-values of each agent is used as the first fitted value of the policy network. This is a relatively mature technique, and this invention will not elaborate on it further. Specifically, the process of using the sum of the Q-values of each agent as the first fitted value of the policy network is implemented through the Value Decomposition Networks (VDN) algorithm.
[0143] (3) Process the training dataset through the target network to obtain the second fitted value of the target network.
[0144] Specifically, the determined training dataset is input into the target network to obtain the fitted value of the target network. For example, the training dataset is input into the target network of the DDPG algorithm to determine the target value of each agent.
[0145] Specifically, the target value for each agent is represented by the following formula:
[0146] y q =r+γ·Q target
[0147] Among them, y q Let represent the target Q value of the q-th agent, r represent the reward corresponding to the training data, and γ represent the coefficient, γ∈(0,1).
[0148] Specifically, the sum of the target values of each agent is used as the second fitted value of the target network. Inputting the determined training dataset into the target network to determine the current Q-value of each agent is a relatively mature technique, and this invention will not elaborate further on it.
[0149] (4) Solve for the error between the first and second fitted values using the mean squared error loss function.
[0150] (5) Determine the policy parameters of the output policy network by using the error and the preset threshold.
[0151] Specifically, the preset threshold refers to the expected reward value. Determining the policy parameters of the output policy network by comparing the error with the preset threshold means that when the error meets the set error threshold, all parameters of the policy network are updated through gradient backpropagation of the deep neural network, and when the reward at the current time meets the expected reward value, all parameters of the updated policy network are used as the policy parameters.
[0152] By implementing this embodiment, a control agent, a first predictive agent, and a second predictive agent are introduced into the integrated energy system. The control agent, the first predictive agent, and the second predictive agent respectively represent the energy storage side, the source side, and the load side in the integrated energy system. Based on the predicted electricity output and the predicted electricity consumption, at least one control parameter of the control agent is determined through a third neural network model. Thus, given the known electricity output on the source side, the energy storage on the energy storage side is adjusted through the determined control parameter, and the load on the load side is also adjusted through the control parameter, thereby achieving energy supply balance in the integrated energy system.
[0153] This embodiment provides a collaborative control device for multiple intelligent agents in an integrated energy system. The integrated energy system includes a first predictive intelligent agent, a second predictive intelligent agent, and a first control intelligent agent. The first control intelligent agent includes at least one adjustable unit, such as... Figure 5 As shown, the device includes: a receiving unit 51, an output unit 52, and a control unit 53, wherein,
[0154] The receiving unit 51 is used for the first control agent to receive a first electrical energy parameter from the first predictive agent and a second electrical energy parameter from the second predictive agent. The first electrical energy parameter is the output electrical parameter predicted by the first neural network model, and the second electrical energy parameter is the consumption electrical parameter predicted by the second neural network model. For details, please refer to the description of step S21 in the above embodiments, which will not be repeated here.
[0155] Output unit 52 is used by the first control agent to process the first and second electrical energy parameters through a third neural network model and output at least one control parameter. For details, please refer to the description of step S22 in the above embodiments, which will not be repeated here.
[0156] The control unit 53 is used by the first control agent to control the operation of at least one adjustable unit through at least one control parameter. For details, please refer to the description of step S23 in the above embodiments, which will not be repeated here.
[0157] One embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions that can execute the collaborative control method of multiple intelligent agents in the integrated energy system described in any of the above method embodiments. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.
[0158] One embodiment of the present invention also provides an electronic device, such as... Figure 6 As shown, Figure 6 This is a schematic diagram of a computer device according to an optional embodiment of the present invention. The computer device may include at least one processor 61, at least one communication interface 62, at least one communication bus 63, and at least one memory 64. The communication interface 62 may include a display screen and a keyboard; optionally, the communication interface 62 may also include a standard wired interface or a wireless interface. The memory 64 may be high-speed RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 64 may also be at least one storage device located remotely from the aforementioned processor 61. The processor 61 may be combined with... Figure 5 The described apparatus has an application program stored in memory 64, and the processor 61 calls the program code stored in memory 64 to perform the steps of the collaborative control method for multiple intelligent agents in the integrated energy system described in any of the above method embodiments.
[0159] The communication bus 63 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 63 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0160] The memory 64 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 64 may also include a combination of the above types of memory.
[0161] The processor 61 can be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP.
[0162] The processor 61 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0163] Optionally, the memory 64 is also used to store program instructions. The processor 61 can invoke the program instructions to implement the cooperative control method for multiple intelligent agents in the integrated energy system described in any embodiment of the present invention.
[0164] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for cooperative control of multiple intelligent agents in an integrated energy system, characterized in that, The integrated energy system includes a first predictive agent, a second predictive agent, and a first control agent, wherein the first control agent includes at least one adjustable unit, and the method includes: The first control agent receives a first power parameter from the first prediction agent and a second power parameter from the second prediction agent. The first power parameter is the output power parameter predicted by the first neural network model, and the second power parameter is the consumption power parameter predicted by the second neural network model. The first control agent processes the first electrical energy parameter and the second electrical energy parameter through a third neural network model and outputs at least one control parameter. The first control agent controls the operation of the at least one adjustable unit through the at least one control parameter; The third neural network model includes a policy network and a target network, wherein the training datasets for the policy network and the target network include training data and rewards corresponding to the control agent. Establish training datasets for the policy network and the target network, including: Obtain the first state of the controlling agent; The first state is processed by the policy network to obtain the action value of the controlling agent; Establish a simulation model corresponding to the adjustable unit; The action value is input into the simulation model to obtain the second state of the control agent; Based on the second state, calculate and control the corresponding reward for the agent; The first state, the action value, and the second state are used as the training data, and the dataset includes the training data and the reward. The calculation of the reward corresponding to the control agent based on the second state includes: Determine the type of the reward; Based on the second state and the type of reward, calculate and control the corresponding reward for the agent; The formula for calculating the reward corresponding to the controlling agent is as follows: in, r This represents the reward at the current moment. r cost Indicates energy cost incentive. r EV This indicates a reward for flexible load control. r com Indicates a comfort reward. ω cost This indicates the weight of the energy cost incentive. ω EV This indicates the weight of the flexible load control reward. ω com This indicates the weight of the comfort reward.
2. The method according to claim 1, characterized in that, Before the first control agent processes the first electrical energy parameter and the second electrical energy parameter through the third neural network model, the method further includes: obtaining the third neural network model. The acquisition of the third neural network model includes: The training dataset is processed by the policy network to obtain the first fitted value of the policy network; The training dataset is processed by the target network to obtain the second fitting value of the target network; The error between the first fitted value and the second fitted value is solved by using the mean squared error loss function; The policy parameters of the policy network are determined by comparing the error with a preset threshold.
3. The method according to claim 1, characterized in that, The calculation of the reward corresponding to the controlling agent includes: The energy cost reward corresponding to the first control agent is determined by the first relation and the second relation. The first relation is: The second relation is: in, r cost Indicates energy cost incentive. i This indicates the number of the first controlling agents. j This represents the number of first-level control agents that can engage in electricity sales. j ∈ i , M n This represents the electricity purchase / sale costs for each primary control agent. W n This represents the electrical energy corresponding to each first control agent. W n A value greater than zero indicates an electricity purchase. W n A value less than zero indicates an electricity sales activity. P pur This indicates the current electricity price. P sell This indicates the current electricity price.
4. The method according to claim 3, characterized in that, The integrated energy system further includes: a second control agent, which includes at least one adjustable unit. Calculating the reward corresponding to the controlling agent further includes: The flexible load control reward corresponding to the second control agent is determined using the third and fourth relational expressions. The third relation is: The fourth relation is: in, r EV This indicates a reward for flexible load control. t Indicates the current moment. t set Indicates the preset time. This indicates the state of charge of the load currently connected to the second control agent. W EV This indicates the charge / discharge level of the second controlling agent. This represents the total capacity of the load connected to the second control agent. This indicates the state of charge of the load connected to the second control agent at the previous moment.
5. The method according to claim 4, characterized in that, The integrated energy system further includes a third control agent, which comprises at least one adjustable unit. Calculating the reward corresponding to the controlling agent further includes: The comfort reward corresponding to the third control agent is determined through the fifth relation. The fifth relation is: in, r com Indicates a comfort reward. T max This indicates the maximum temperature value within the comfortable temperature range. T min This represents the minimum temperature value within the comfortable temperature range. Indicates the first k The set temperature value of the third control agent. l This indicates the number of third-party control agents.
6. A collaborative control device for multiple intelligent agents in an integrated energy system, characterized in that, The integrated energy system includes a first predictive agent, a second predictive agent, and a first control agent, wherein the first control agent includes at least one adjustable unit, and the device includes: The receiving unit is configured to receive, respectively, a first electrical energy parameter from the first predictive agent and a second electrical energy parameter from the second predictive agent, wherein the first electrical energy parameter is the output electrical parameter predicted by the first neural network model and the second electrical energy parameter is the consumption electrical parameter predicted by the second neural network model. The output unit is used by the first control agent to process the first electrical energy parameter and the second electrical energy parameter through the third neural network model and output at least one control parameter. A control unit is used by the first control agent to control the operation of the at least one adjustable unit through the at least one control parameter; The third neural network model includes a policy network and a target network, wherein the training datasets for the policy network and the target network include training data and rewards corresponding to the control agent. Establish training datasets for the policy network and the target network, including: Obtain the first state of the controlling agent; The first state is processed by the policy network to obtain the action value of the controlling agent; Establish a simulation model corresponding to the adjustable unit; The action value is input into the simulation model to obtain the second state of the control agent; Based on the second state, calculate and control the corresponding reward for the agent; The first state, the action value, and the second state are used as the training data, and the dataset includes the training data and the reward. The calculation of the reward corresponding to the control agent based on the second state includes: Determine the type of the reward; Based on the second state and the type of reward, calculate and control the corresponding reward for the agent; The formula for calculating the reward corresponding to the controlling agent is as follows: in, r This represents the reward at the current moment. r cost Indicates energy cost incentive. r EV This indicates a reward for flexible load control. r com Indicates a comfort reward. ω cost This indicates the weight of the energy cost incentive. ω EV This indicates the weight of the flexible load control reward. ω com This indicates the weight of the comfort reward.
7. An electronic device, characterized in that, include: At least one processor; and a memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor. When the instructions are executed by the at least one processor, the electronic device performs the collaborative control method of multiple intelligent agents in an integrated energy system as described in any one of claims 1 to 5.
8. An integrated energy system, characterized in that, The integrated energy system includes a first predictive agent, a second predictive agent, and a first control agent. The first control agent includes at least one adjustable unit. The first predictive agent is used to collect first meteorological data, input the first meteorological data into the first neural network model, and obtain the first electrical energy parameter, which is the predicted output electrical parameter; The second predictive agent is used to collect second meteorological data, input the second meteorological data into the second neural network model, and obtain the second electrical energy parameter, which is the predicted electrical consumption parameter; The first control agent receives a first power parameter from the first prediction agent and a second power parameter from the second prediction agent. The first control agent processes the first and second electrical energy parameters through a third neural network model and outputs at least one control parameter. The first control agent controls the operation of the at least one adjustable unit through the at least one control parameter; The third neural network model includes a policy network and a target network, wherein the training datasets for the policy network and the target network include training data and rewards corresponding to the control agent. Establish training datasets for the policy network and the target network, including: Obtain the first state of the controlling agent; The first state is processed by the policy network to obtain the action value of the controlling agent; Establish a simulation model corresponding to the adjustable unit; The action value is input into the simulation model to obtain the second state of the control agent; Based on the second state, calculate and control the corresponding reward for the agent; The first state, the action value, and the second state are used as the training data, and the dataset includes the training data and the reward. The calculation of the reward corresponding to the control agent based on the second state includes: Determine the type of the reward; Based on the second state and the type of reward, calculate and control the corresponding reward for the agent; The formula for calculating the reward corresponding to the controlling agent is as follows: in, r This represents the reward at the current moment. r cost Indicates energy cost incentive. r EV This indicates a reward for flexible load control. r com Indicates a comfort reward. ω cost This indicates the weight of the energy cost incentive. ω EV This indicates the weight of the flexible load control reward. ω com This indicates the weight of the comfort reward.
Citation Information
Patent Citations
Micro-grid operational control method based on multiple agents
CN103730891A
Microgrid energy storage scheduling method, device and equipment based on deep reinforcement learning
CN112529727A