Method and system for resilient agent energy control
A decentralized reinforcement learning method using a flexibility signal and adaptive policy matrix addresses the challenges of managing multiple buildings with diverse control strategies, ensuring efficient, adaptive, and privacy-focused energy management.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NIK VAHID MOUSSAVI
- Filing Date
- 2025-11-12
- Publication Date
- 2026-05-21
AI Technical Summary
Existing energy management systems face challenges in managing large groups of buildings with multiple control strategies, requiring high computational power, data sharing, and centralized decision-making, which can compromise privacy and adaptability to environmental changes.
A decentralized reinforcement learning-based method that uses a flexibility signal from a central server to control energy performance of agents without data sharing, employing a model-free RL engine with a multidimensional policy matrix and adaptive updating to enhance scalability and resilience.
Enables secure, efficient, and adaptive energy management of large urban areas with reduced computational load, ensuring privacy and rapid adaptation to environmental changes, while maintaining system stability and user comfort.
Smart Images

Figure 00000036_0000 
Figure 00000037_0000 
Figure 00000038_0000
Abstract
Description
[0001] METHOD AND SYSTEM FOR RESILIENT AGENT ENERGY CONTROL FIELD OF INVENTION
[0002] The invention relates to a system for controlling energy performance of an agent within an energy network in which energy is distributed to a plurality of agents. The energy network is operated by a distribution system operator (DSO) for distribution of energy, such as electric energy or heat energy, to a plurality of agents.
[0003] BACKGROUND ART
[0004] Synchronized energy management (EM) of buildings and energy systems in urban areas can become computationally complex and practically expensive. Multiple approaches are previously available for demand response (DR) and Demand Side Management (DSM).
[0005] Energy management is needed for increasing the flexibility of energy systems, resulting in a higher reliability and carbon reduction of energy solutions without heavy network investments. The implementation of EM in urban areas may have been lagging behind due to increased complexity of the system operation and the need for expensive ICT solutions and implying privacy and security risks.
[0006] A big challenge exists in finding optimum building controls for a large group of buildings connected to the energy grid and influencing each other. Buildings usually have multiple control strategies, so reaching an optimum control can become computationally demanding. Moreover, there exists several other influencing factors and indicators, such as user behavior / comfort and appliances, which all can affect the energy performance of buildings. Considering multiple factors and reaching an optimum control strategy can become very challenging in simulating and controlling buildings in connection to the energy grid.
[0007] Moreover, there is a need for lighter algorithms to be implemented in energy management systems, enabling a large number of agents (with multiple control strategies) to communicate and collaborate in a complex environment. Enhanced, meanwhile simple to implement, approaches are needed when the number of buildings and control strategies increase, especially when preparing for future extreme and unprecedented events.
[0008] To this end, implementation of Reinforcement Learning (RL) has shown promising results, especially the model-free RL approaches (e.g. Q-learning). RL-based methods have shown substantial potential in resolving increasing complexities within the energy domain, considering both supply and demand.
[0009] A major limitation in the current state of the art in implementing RL is the oversimplification of energy problems. There is a large need to develop RL-based methods for real building applications to accelerate training and enhance control robustness, especially with larger participation of multiple agents with different priorities. Robust and reliable RL policies are needed to address environmental shifts and mismatched configurations. RL algorithms with reduced variance are needed to control multi-agent systems in non-stationary environments.
[0010] Addressing this gap is important to effectively adapt to the escalating frequency and intensity of climate variations. Although implementation of RL in energy control and optimization is not new, there is still no approach to work easily (with low cost and computational power) and without data sharing when the number of buildings / agents increase. The available methods mostly rely on transferring data outside the building boundaries (cloud computation), where optimization and decision making mostly happens on servers and transferred through ICT means to buildings / agents. In other words, a big part of decision making happens at a central level outside buildings.
[0011] There is a need for more secure and decentralized decision-making approaches that guarantee user’s privacy. Moreover, coping with shocks and climate variations requires agile and less centralized systems. This calls for innovative RL-based solutions that are currently absent in the available technology.
[0012] Patent publication W02019063079A1 discloses a system, device and method for energy and comfort optimization in a Building Automation Environment. The method includes receiving environment data associated with the building automation environment. A building model is generated for the building automation environment. The building model is represented by a set of states comprising energy profiles and comfort profiles. Reward vectors for the set of states with the energy profiles and the set of states with the comfort profiles are determined. The reward vectors are determined based on probabilities of transition from a current state to remaining states of the set of states. An optimization policy for energy and comfort is determined based on the reward vectors. An action is performed based on the optimization policy whereby the building automation environment transitions to a new state in the set of states. Accordingly, the energy and comfort are optimized by iteratively determining the reward vectors and the optimization policy for the new state.
[0013] However, the system according to W02019063079A1 needs a model of the building, which is a pre-trained simulation model using historical environment data and predicted environment data and including multi-layer neural network. The system is mainly developed to optimize the performance of a single building / agent. It may be possible to implement the system in several buildings, but it becomes computationally expensive and needs to have the models of buildings.
[0014] TECHNICAL PROBLEM
[0015] There is a need for a technology that does not need to know about building characteristics. Such new technology may comprise information about control options and energy / comfort conditions without the need to understand the building. The same control algorithm can be arranged in any building without taking any effort of modeling buildings. It also does not need to share data with any server or similar, it just receives data (e.g. one signal) from the network. Such simplicity enables to have many controllers / actuators connected while still keeping the computational load very low. It also enables to manage large urban areas without the need for high computational power or data servers.
[0016] SUMMARY OF THE INVENTION
[0017] Accordingly, an object of the present invention is to mitigate, alleviate or eliminate one of more of the above-mentioned shortcomings and other deficiencies and disadvantages, singly or in any combination.
[0018] According to a first aspect, a computer-implemented method for controlling energy performance of an agent among a plurality of agents within an energy network in which energy is distributed to the plurality of agents is provided. The method being performed by an optimization and control unit associated with the agent.
[0019] The method comprises receiving a flexibility signal from a central server system, external to and operating above the plurality of agents. The flexibility signal being indicative of the total energy load of the plurality of agents in relation to a predefined reference value and reflecting a current or upcoming environmental condition. The flexibility signal being a digital signal.
[0020] The method further comprises operating a model-free reinforcement learning, RL, engine within the optimization and control unit to select a control action based on the received flexibility signal, a current time, and a multidimensional policy matrix, wherein the multidimensional policy matrix comprises a dimension representing the flexibility signal, a dimension representing time, and a dimension representing a set of control actions. The time may e.g. be a date, an hour of a date, and / or part of an hour of the date. A control action may e.g. be an action that controls a temperature setpoint, a ventilation rate and / or heating / cooling power.
[0021] The method enables decentralized control of agents without data sharing between the agents or a central server. Within an energy network, an agent is typically defined as an energy user. However, an agent may be as energy user / producer (prosumer), an energy storage, and an energy producer. Hence, an agent may be an energy user, an energy producer, an energy prosumer, or an energy storage. An agent can be any combination of these. The energy network comprises a plurality of agents. These agents can vary in scale, from individual buildings or groups of buildings to parts of the energy network itself. Each agent operates as an independent unit, making decisions to control its own energy performance without sharing data or being aware of other agents in the energy network. They are equipped with an optimization and control unit that uses reinforcement learning, RL, to manage their energy use. This decentralized approach allows for a secure and private system, as decision-making happens locally at the respective agent. Behavior at an agent is influenced by the flexibility signal centrally generated at a central server system external to and operating above the plurality of agents (e.g., from a distribution system operator, DSO). Control actions at the agent, such as adjusting heating or cooling setpoints and / or adjusting electricity usage, are its way of responding to this flexibility signal and adapting to environmental conditions. The flexibility can be any signal in the case of need, representing grid conditions, price, weather events, shocks, etc.
[0022] The method provides several potential technical advantages for controlling energy performance in an energy network. It addresses common challenges in decentralized, multi-agent systems, such as computational complexity, data privacy, and adaptability to extreme environmental events. The method enables a network of agents to operate independently, with each agent making decisions based on a single, universal flexibility signal from a central server system. This eliminates the need for agents to share data with each other or with a central server for optimization. This decentralized approach significantly enhances user privacy and data security, as decision-making happens locally at the agent level. The reliance on a single, global flexibility signal allows for a minimum level of data transfer between agents and the energy provider. This removes the need for high computational power at the central level, making the method scalable to manage large urban areas or a vast number of agents without requiring expensive ICT solutions or high-power data servers. Hence, the optimization and control unit within each agent may be a lightweight algorithm designed for low computational power and enabling management of a large number of agents.
[0023] The method may also be used for a plurality of energy networks, e.g. managing multiple microgrids connected to a central grid. In such case the different energy networks (or microgrids) are to be seen as an agent.
[0024] The method may further comprise assigning a reward to possible control actions in the set of control actions using a reward mechanism, the reward mechanism being a conditional selection system which in addition to a reward of 1 for a control action achieving a goal and a reward of 0 for other control actions comprises at least one weighted reward control action which relates to a control action fulfilling a primary goal while compromising a secondary goal. The reward mechanism may hence comprise the rewards 0, 0.5, and 1. According to another example, the reward mechanism may comprise the rewards 0, 0.3, 0.5, 0.8, and 1. These are just some examples of possible rewards. An example of a primary goal is energy reduction. An example of a secondary goal is indoor comfort. The reward function may comprise multiple secondary goals, for example, a combination of energy saving, indoor comfort, CO2 reduction, each with its assigned reward.
[0025] The method may further comprise conditionally selecting a control action that corresponds to a weighted reward control action when the flexibility signal indicates extreme conditions on the energy network, thereby controlling energy performance of the agent within the energy network by reducing the overall energy demand of the agent and contributing to a system-wide response to extreme conditions.
[0026] The provided method presents a weighted reward mechanism. By assigning a weighted reward (e.g., 0.5) to a control action that fulfills a primary goal (e.g., energy reduction) while compromising a secondary one (e.g., indoor comfort), the system gains a higher degree of agility. This conditional selection enables agents to respond faster to "high pressure" or extreme conditions, such as heatwaves or cold waves, by selecting a less-than-perfect action that helps the entire network pass the extreme condition without risking an energy shortage. The use of a multidimensional policy matrix combined with the ability to "bundle control actions over time" accelerates the reinforcement learning process. By gathering control actions for similar conditions (e.g., specific hours of the day over an entire season), the system can evaluate a large number of control actions over a short period to quickly find and store optimal ones. This contrasts with traditional RL methods that may take a long time to reach an optimal policy.
[0027] The method may further comprise updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions. Hence, the method may be configured with an adaptive mechanism to automatically update the policy matrix to suit changing environmental conditions. This is achieved by removing "weakest control actions" and replacing them with potentially better ones, ensuring the method remains prepared for unprecedented situations. Hence, it is enabled for the method to adapt and remain effective in non-stationary or changing environments without human intervention or centralized control. By periodically or conditionally removing the least effective actions, the system can focus its learning on more promising control strategies. This acts as a form of pruning or curriculum learning, allowing the RL engine to find an optimal policy faster than if it had to explore a vast and unchanging set of suboptimal actions. When the method encounters a new or "unprecedented" environmental condition, such as a more extreme heatwave than previously experienced, the existing policy may contain outdated or ineffective actions. The adaptive update mechanism ensures that the method is not locked into these suboptimal strategies. Instead, it can automatically adjust its policy to better fit the new reality, improving its resilience and performance over time. This is an advantage, especially in the context of climate change and extreme weather events. The method's ability to automatically update the policy matrix removes the need for a central controller to manually intervene and retrain the method. This not only reduces operational costs but also allows the method to remain highly efficient and responsive, as it continuously refines its decision-making capabilities based on real-world performance. The adaptation happens at the agent level, meaning it is decentralized and scalable.
[0028] Updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions may be performed periodically. By periodically is in this context meant within a certain time period, e.g. every week, every second week, once a months, etc.
[0029] Updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions may be performed upon a condition occur. Examples of such conditions are reduced energy saving or sudden shocks such as extreme climate variations.
[0030] The weakest control actions may be identified as those with the lowest assigned reward. By linking "weakest" directly to the lowest assigned reward, the method has a quantifiable and objective basis for its self-correction. The reward function, which accounts for factors like energy saving and indoor comfort, provides a direct measure of an effectiveness of a control action. This ensures that the adaptation is grounded in real-world performance data. The method facilitates a form of "learning from mistakes." Instead of holding onto a vast library of control actions, many of which are suboptimal, the method can systematically prune its knowledge base. This reduces the computational overhead of having to evaluate a large set of control actions at each time step. By replacing the least-rewarded actions, the system becomes more focused and efficient over time, as it continually refines its policy with control actions that have proven to be more successful. This approach allows the method to dynamically tune its performance based on its defined goals. For example, if a weighted reward of 0.5 is assigned to a control action that prioritize energy savings over comfort, the system will identify actions that do not meet this goal as "weak." By replacing these with new, more effective control actions, the system automatically adapts to better achieve its primary objective, especially during periods of high network pressure or "extreme conditions". Conditionally selecting a control action may comprise selecting the control actions only from a predefined number of nearest control actions relative to a current control action. This enables us to ensure a smooth transition in the control system and prevent sudden jumps. By preventing sudden or drastic "jumps" between control actions, the method ensures a smooth transition in the control system. Sudden changes may cause indoor discomfort and accelerate wear and tear of mechanical and electrical systems. Limiting the selection to the nearest actions mitigates this risk and promotes more predictable and reliable performance. Many actuators and controllers, like those for heating, ventilation, and air conditioning (HVAC) systems, are not designed for frequent or sudden changes. Limiting the selection to nearby actions reduces the mechanical stress on these devices, thereby extending their operational lifespan and decreasing maintenance costs. The method accounts for the physical nature of the controlled systems.
[0031] Restricting the search space for the next optimal action can actually make the learning process more efficient. Instead of having to evaluate all possible control actions at every time step, the RL engine only needs to consider a smaller, more relevant subset of control actions. This can speed up the decision-making process and reduce computational load. In applications like building energy management, user comfort is a critical factor. Sudden, large changes in temperature or other indoor climate parameters can cause discomfort. By enabling smooth transitions, the method ensures that the system's control actions do not negatively impact the occupants' experience. The "tagging system" is flexible, meaning the selection criteria can be different for various controllers and agents within the same network. This allows for a high degree of customization, where system operators can set different transition limits based on the specific type of equipment or the desired level of control, enhancing the freedom of the system's design and selection criteria.
[0032] The flexibility signal may be a price signal. By using a price signal that reflects the cost or value of energy on the grid, the method directly encourages agents to reduce their energy demand during periods of high prices (high load on the network) and increase it (in the case of need) during periods of low prices (high supply). This aligns individual agent behavior with the overall network's needs in a way that is both effective and transparent.
[0033] The method may further comprise generating the flexibility signal at the central server system by. Generating the flexibility signal may comprises: defining a range of values that indicate the need for flexibility; comparing the current energy load of the plurality of agents to a set of predefined reference values for a specific time or season; and assigning a digital signal from the defined range based on the comparison. By defining a simple, limited range of digital values (e.g., 0-5), the system creates a universal "language" for communication within the energy network. This standardization makes it easy for diverse agents to understand and respond to the signal, regardless of their specific hardware or control logic. It eliminates the complexity of having to interpret a continuous flow of data, which can fluctuate rapidly. The method of comparing the current energy load to predefined reference values is a lightweight algorithm. This comparative process requires minimal computational power, enabling it to be implemented easily at a central level without the need for large, expensive data servers. This makes the system highly scalable and suitable for managing a large number of agents or urban areas. The digital signal is generated to reflect a specific, easily understood condition: the balance between energy demand and supply. For example, a signal of 'O' means there's no need for flexibility, while a '5' indicates a high-pressure situation requiring maximum flexibility, such as a large reduction in energy demand. This transparency allows for targeted demand-side management, ensuring that agents only take action when it is truly needed to prevent grid stress or energy shortages. The flexibility signal can be specifically shaped around the purpose of enhancing climate resilience. By using historical data from a typical season as a reference, the signal can be configured to respond effectively when current conditions deviate from the norm, such as during a heatwave or a cold wave. This makes the system proactive in handling extreme and unprecedented events, as the signals are designed to compel agents to adapt to new environmental conditions.
[0034] Generating the flexibility signal may comprise: obtaining real-time data related to at least one environmental factor, such as solar irradiance, wind speed, or ambient temperature, from one or more sensors; analyzing the real-time data to determine a current state of the environmental condition; and generating the flexibility signal based on the current state, wherein the signal represents the immediate availability or constraint of a resource influenced by the environmental condition, such as the generation capacity of a renewable energy source. By such generation of the flexibility signal, the method may directly link the energy demand-side management to realtime supply-side conditions. By using sensor data for factors like solar irradiance or wind speed, the central server system can immediately assess the current generation capacity of renewable energy sources. This allows the flexibility signal to be dynamically adjusted based on the immediate availability of clean energy, enabling agents to instantly increase or decrease their energy consumption to match the current supply. This real-time synchronization is a significant advantage over static, time-based signals. By reflecting the immediate availability of a resource, the flexibility signal helps the network maintain stability. For example, if wind speed suddenly drops and a wind farm's output decreases, a real-time signal can be sent to all agents to reduce their energy demand, thereby preventing a potential imbalance between supply and demand that could lead to brownouts or blackouts. This proactive, data-driven approach allows for more robust grid management and prevents costly and disruptive outages. The method may directly facilitate the integration of intermittent renewable energy sources into the grid. Historically, a major challenge with renewables like solar and wind has been their unpredictable nature. By providing a mechanism for the network to respond to real-time changes in renewable energy generation, the method maximizes the use of clean energy when it's available and minimizes reliance on fossil fuels. This leads to higher overall energy efficiency and a greater degree of sustainability for the entire network.
[0035] The method may further comprise receiving a predictive environmental signal indicative of at least one future environmental factor from an external source and modifying the selection of the control action based on the predictive environmental signal. Modifying the selection of the control action based on the predictive environmental signal may be made by implementing precooling or pre-heating strategies. By this the agent's energy demand in preparation for the upcoming environmental condition may be adjusted. The at least one future environmental factor may comprise one or more of predicted solar irradiance, forecasted temperature, forecasted wind speed, or anticipated cloud cover. The external source may be a weather forecasting service.
[0036] This process, which may be referred to as predictive reinforcement learning, allows the agents to behave proactively. By knowing one or more future environmental factors in advance, the agent can initiate energy demand adjustments before the event occurs. For instance, a building can slightly pre-cool its space during a period of low electricity price or low grid demand, storing "coolness" in its thermal mass, to reduce the need for high demand cooling later when a heatwave is forecasted. The agent performs this pre-conditioning autonomously based on the received predictive environmental signal, demonstrating the system's ability to learn and exploit the thermal capacity of buildings without needing to know anything about them (i.e., it's model-free RL). The predictive functionality may improve the stability and user-friendliness of the control system. The predictive capability is particularly beneficial for coping with climate shocks and variations. The agent is informed about the upcoming conditions (e.g., a cold wave) and is driven to prepare for the new environmental conditions. Simulation results show that using the predictive feature can result in saving more energy during extreme conditions. For instance, during cold wave periods, the predictive feature kept the indoor temperature at the lower bound around 25% more than the 'No forecast' case resulting is less energy demand from the grid during extreme climate events. The use of a separate predictive environmental signal allows the system to remain effective across geographically diverse networks. Separating the predictive environmental signal from the main flexibility signal allows a large grid to factor in local weather conditions that might be different at different parts of the grid. This enables a more tailored and accurate response from the local agents. Since the agent autonomously modifies its action selection based on the predictive environmental signal, the centralized control system doesn't need to perform complex, localized climate calculations, keeping the overall architecture lightweight and simple to implement.
[0037] The flexibility signal may be a single global signal transferred to all agents in the network per time step.
[0038] The flexibility signal may be treated in different ways for agents with different natures. For example, if we have a network of 10 residential buildings and one hospital, the residential buildings will decrease their demand by signal 5 as much as they can, but hospital does not need to do it since they are considered as sensitive buildings. This is thanks to local adjustments of the algorithm per agent.
[0039] Hence, a novel energy management approach based on integrating core elements of reinforcement learning, RL, is provided and empowered by adaptive reinforcement learning. This will herein be referred to as Adaptive Reinforcement Learning for Energy Management, ARLEM. The technology significantly focuses on this pivotal aspect, presenting a novel approach to meet these pressing challenges. In ARLEM, the RL engine, implemented at the agent level, may be made to have adaptive and predictive characteristics, such energy management will herein be referred to as Predictive and Adaptive Reinforcement Learning for Energy Management, PARLEM. Its adaptive characteristic enables agents to update their policies (set of control actions) in relation to the variations of their surrounding environment. The predictive characteristic enables agents to prepare for the upcoming conditions (of the environment) in advance. Each of these characteristics, as well as their combination, make the whole system agile and more efficient.
[0040] One of the novel contributions of ARLEM (compared to the available RL-based methods) is the logic for creating its policy, which enables the technology to quickly reach optimum solutions and adapt to the environmental variations. It also enables an ensemble of agents to act synergistically without knowing about each other or communicating any data / information.
[0041] According to a second aspect, an optimization and control unit associated with an agent among a plurality of agents within an energy network in which energy is distributed to the plurality of agents is provided. The optimization and control unit is configured to control energy performance of the associated agent. The optimization and control unit comprising processing circuitry for carrying out the method according to the first aspect. The above-mentioned features and potential advantages of the method according to the first aspect, when applicable, apply to this second aspect as well. In order to avoid undue repetition, reference is made to the above.
[0042] According to a third aspect, a computer program is provided. The computer program comprising instructions which, when the program is executed by optimization and control unit having processing capabilities, cause the optimization and control unit to perform the method according to the first aspect.
[0043] According to a fourth aspect, a computer-readable storage medium having stored thereon a computer program according to the third aspect is provided.
[0044] The above-mentioned features and potential advantages of the method according to the first aspect, when applicable, apply to the third and fourth aspects as well. In order to avoid undue repetition, reference is made to the above.
[0045] According to a fifth aspect, a system for controlling energy performance of a plurality of agents in an energy network is provided. The system comprises: a central server system configured to generate a flexibility signal, the flexibility signal being a digital signal indicative of the total energy load of the plurality of agents in relation to a predefined reference value and reflecting a current or upcoming environmental condition; and a plurality of optimization and control units, each being assigned to a separate one of the plurality of agents. Each optimization and control unit is configured to: receive the flexibility signal, operate independently without awareness of other optimization and control units, operate a model-free reinforcement learning, RL, engine to select a control action based on the received flexibility signal, a current time, and a multidimensional policy matrix, wherein the multidimensional policy matrix comprises a dimension representing the flexibility signal, a dimension representing time, and a dimension representing a set of control actions. There is further a potential to add more dimensions, each controller having a separate vector for itself.
[0046] Each optimization and control unit may further be configured to assign a reward to possible control actions in the set of control actions using a reward mechanism, the reward mechanism being a conditional selection system which in addition to a reward of 1 for a control action achieving a goal and a reward of 0 for other control actions comprises at least one weighted reward control action which relates to a control action fulfilling a primary goal while compromising a secondary goal, and conditionally select an action corresponding to a weighted reward control action when the flexibility signal indicates extreme conditions within the energy network, thereby controlling energy performance of the agent by reducing its overall energy demand of the agent and contributing to a system-wide response to extreme conditions. Each optimization and control unit may further be configured to update the multidimensional policy matrix by replacing a subset of weakest actions with new actions.
[0047] Updating of the multidimensional policy matrix may be performed periodically. Alternatively, or in combination, updating of the multidimensional policy matrix may be performed when a condition occurs.
[0048] Each optimization and control unit may further be configured to conditionally select a control action by selecting the control action only from a predefined number of nearest control actions relative to a current control action.
[0049] Generating the flexibility signal at the central server system may comprise: define a range of values that indicate the need for flexibility; compare the current energy load of the plurality of agents to a set of predefined reference values for a specific time or season; and assign a digital signal from the defined range based on the comparison.
[0050] The central server system may be configured to generate the flexibility signal by: obtaining real-time data related to at least one environmental factor, such as solar irradiance, wind speed, or ambient temperature, from one or more sensors; analyzing the real-time data to determine a current state of the environmental condition; and generating the flexibility signal based on the current state, wherein the signal represents the immediate availability or constraint of a resource influenced by the environmental condition.
[0051] Each optimization and control unit may further be configured to: receive a predictive environmental signal indicative of at least one future environmental factor, such as predicted solar irradiance, forecasted temperature, forecasted wind speed, or anticipated cloud cover, from an external source and modify the selection of the control action based on the predictive environmental signal.
[0052] The flexibility signal may be a single global signal transferred to all agents in the network per time step.
[0053] Each of the plurality of agents may be a building, a part of a building or part of the energy network.
[0054] A same flexibility signal may be provided to all optimization and control units for providing demand-side management.
[0055] In an further aspect, there is provided a system for controlling behavior of multiple agents comprising: an energy network with multiple agents, working as energy user, producer or storage, at different scales, e.g. building, groups of buildings, or energy grid / network, operated by a control or distribution system, for example, building management system (BMS) or distribution system operator (DSO), for controlling the energy performance of agents and / or distribution of energy, such as electric energy or heat energy; a network of agents operating independently substantially without awareness of each other and without sharing data with another agent; an optimization and control unit, being implemented in each agent and working based on the core concepts of reinforcement learning implementing model-free reinforcement learning (RL); a mechanism to generate a universal signal, such as a flexibility signal or a price signal, to represent environmental conditions, with or without predictive features, and providing said universal signal to the agents in the network, the universal signal being emitted from one level or more above the agent level, reflecting upon the total performance of the agents or conditions of the environment; a multidimensional policy matrix per agent, each dimension representing a key characteristic of the environment and possible control actions; a reward mechanism for conditional selection and higher agility of the system, through weighting the reward signals in several steps, for example 0, 0.1, 0.3, 0.5, 0.7, 0.9 and 1.
[0056] In an embodiment, the system may further comprise: an action selection (tagging) mechanism, to enable smooth transition in the control system, such as picking the next action only from nearest actions or two nearest actions. The system may further comprise: an adaptive mechanism updating the multidimensional policy matrix according to defined settings, such as a time span to update and the number of actions. The system may further comprise: an adaptive mechanism based on predictive reinforcement learning for updating the multidimensional policy matrix accordingly by picking suitable actions according to the upcoming conditions over the next few hours. The system may further comprise: an adaptive mechanism based on combining predictive and adaptive reinforcement learning, updating the multidimensional policy matrix through updating policies and accounting for the upcoming conditions. The system may further comprise: a flexible reward function that can account for multiple influencing factors, such as energy saving and indoor comfort. The system may further comprise: a mechanism to induce certain random selection of actions, such as 10% randomness, enabling the agents to explore for new control actions and replace the satisfactory ones with the weakest actions in their library of actions.
[0057] In another embodiment, the agents may be buildings or parts of buildings or parts of the energy network. The same universal signal may be provided to all agents for providing demandside management.
[0058] BRIEF DESCRIPTION OF DRAWINGS
[0059] Further objects, features and advantages will become apparent from detailed description of embodiments with reference to the drawings, in which: Fig. l is a schematic representation of the multidimensional policy matrix in ARLEM, drawn for a 3D sample policy.
[0060] Fig. 2 is a schematic representation showing the status of the system during the considered time span being divided into days and specific hours of days bundled together.
[0061] Fig. 3 is a schematic representation showing an example of flexibility signal generation. Fig. 4 is a schematic representation showing a schematic presentation of ARLEM.
[0062] Fig. 5 is a schematic representation showing illustration of ARLEM at the urban area, cluster node, and edge node (top) showing the Flexibility signal from DSO / TSO transferred to cluster node (green nodes).
[0063] Fig. 6 is a schematic representation showing the adaptive RL engine that enables ARLEM to update the policy over time.
[0064] Fig. 7 is a schematic representation showing energy demand during summer with (HW) and without (TDY) heatwave in Madrid.
[0065] Fig. 8 is a schematic representation showing cumulative cooling demand and degree days over summer in Madrid.
[0066] Fig. 9 is a schematic representation showing energy demand during summer with (HW) and without (TDY) heatwave for different policy update and control limit strategies of ARLEM.
[0067] Fig. 10 is a schematic representation showing distribution of the indoor temperature with ARLEM with three control limits and without any policy update.
[0068] Fig. 11 is a schematic representation showing distribution of the indoor temperature with ARLEM with 10P policy update and three control limits.
[0069] Fig. 12 is a schematic representation showing distribution of the indoor temperature with ARLEM with 20P policy update and three control limits.
[0070] Fig. 13 is a schematic representation showing energy demand during winter with (CW) and without (TDY) cold wave in Stockholm.
[0071] Fig. 14 is a schematic representation showing cumulative heating demand and degree days over winter in Stockholm for different modes of running ARLEM.
[0072] Fig. 15 is a schematic representation showing energy demand during winter with (CW) and without (TDY) cold wave for different policy update and control limit strategies of ARLEM.
[0073] Fig. 16 is a schematic representation showing distribution of the indoor temperature with ARLEM with three control limits and without any policy update.
[0074] Fig. 17 is a schematic representation showing distribution of the indoor temperature with ARLEM with 10P policy update and three control limits. Fig. 18 is a schematic representation showing distribution of the indoor temperature with ARLEM with 20P policy update and three control limits.
[0075] Fig. 19 is a schematic representation showing temporal distribution of indoor setpoint temperatures when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0076] Fig. 20 is a schematic representation showing temporal distribution of indoor temperatures when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0077] Fig. 21 is a schematic representation showing temporal distribution of energy demands of buildings when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0078] Fig. 22 is a schematic representation showing distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0079] Fig. 23 is a schematic representation showing distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0080] Fig. 24 is a schematic representation showing distribution of the indoor temperature in all the buildings during winter when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0081] Fig. 25 is a schematic representation showing distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with no selection criteria (Jump) runs with predictive function for four different forecast spans.
[0082] Fig. 26 is a schematic representation showing distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with TwoStep selection criteria runs with predictive function for four different forecast spans.
[0083] Fig. 27 is a schematic representation showing distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with OneStep selection criteria runs with predictive function for four different forecast spans.
[0084] Fig. 28 is a schematic representation showing distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with no selection criteria (Jump) runs with predictive function for four different forecast spans. Fig. 29 is a schematic representation showing distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with TwoStep selection criteria runs with predictive function for four different forecast spans.
[0085] Fig. 30 is a schematic representation showing distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with OneStep selection criteria runs with predictive function for four different forecast spans.
[0086] DETAILED DESCRIPTION OF EMBODIMENTS
[0087] Below, several embodiments of the invention will be described for enabling a skilled person to perform the invention. These embodiments are given for illustrating the invention and do not limit the scope of the invention, which is solely defined by the appended patent claims.
[0088] General overview of the developed system
[0089] The system according to embodiments can be used in a grid providing electric energy or heating / cooling to buildings. The energy is generated by an energy distributor or supplier and the distribution is administered by a distribution system operator (DSO). The energy is provided to agents or energy users, such as buildings. The energy supplier or DSO generates a universal signal, such as a flexibility signal, that reflects upon the current and / or upcoming conditions and the balance between energy demand and supply. The flexibility signal is generated considering a group (or groups) of buildings (energy users) while buildings are supposed to react to the signal and adapt their demand based on the received signal.
[0090] Flexibility signal and synchronized behavior in the grid
[0091] The need for flexibility is usually transferred from the energy provider to the consumer through the energy network, which is called flexibility signal. In many cases, the price signal may be used as the flexibility signal, affected by the market and bidding strategies that govern the energy grid / market. In ARLEM, the flexibility signal can be any meaningful commutation signal between buildings and the energy grid such as the price signal or a digital signal. Both approaches have been tested and in the present disclosure results for the digital signal are presented. The digital signal is set as a number between 0 and 5, which 0 means there is no need for flexibility and 5 asks for the maximum possible flexibility, which is usually interpreted as lowering the energy demand in buildings as much as possible. The flexibility signal or data signal may alternatively be any number, for example between 0 and 100 or whatever is appropriate in the specific situation. Hence, the digital signal represents data as a series of discrete, distinct values, to convey information in distinct steps rather than a continuous flow. Unlike an analog signal, which has an infinite number of possible values, a digital signal can only take on a finite number of values at any given time, allowing for more reliable transmission of data.
[0092] The flexibility signal is generated at one level above buildings or higher, e.g. generated by the energy supplier or distribution system operator (DSO). It reflects upon the current and / or upcoming conditions and the balance between energy demand and supply. The flexibility signal is generated considering a group (or groups) of buildings (energy users) while buildings are supposed to react to the signal and adapt their demand based on the received signal.
[0093] Generating the flexibility signal
[0094] Different strategies / logics can be adopted to generate the flexibility signal, depending on factors such as regional and energy market policies, available energy sources, climate etc. Both price signals and digital signals have been tried, wherein the former consider more factors into decision making as is the nature of setting the energy price. Results of using the digital signal are further discussed in this disclosure as is easier to explain and demonstrate.
[0095] The logic that has been developed for generating the flexibility signal is very much shaped around the initial purpose of the work which is enhancing the climate resilience of the energy systems. In this regard, the energy demand of the considered building stock in a season (e.g. winter) during typical weather conditions is used as a reference. The span between the average and maximum energy demand is divided into four sections (equal or non-equal sections) and five values (including average and maximum energy demand) are used to generate the flexibility signal in a comparative way. These values have been calculated for the specific hours in the season, for example, the average and maximum energy demand at 20:00 in winter (considering all the winter days). In a winter day, if the energy demand at 20:00 is equal or below the average value at 20:00, the flexibility signal will be 0, if it is above average and less than or equal the next value, the signal will be 1 and so on. For the last signal, 5, we may multiply the maximum typical energy demand by e.g. 0.9 to act more conservatively. The time step for transferring the flexibility signal depends on the case, needs and energy network. In the present case study, there has been adopted a 15-min time step. The time step may be different, for example each minute or each second hour. It does not need to be fixed but can for example vary from 30 minutes for the first six hours of a day, to 20 minutes for the next 12 hours and then back to 30 minutes for the last six hours of the day. Communicating the flexibility signal
[0096] The approach of sending the flexibility signal can differ depending on the intelligence and control level at the energy supply or cluster side (e.g. DSO). Higher control access often corresponds to increased data transfer while compromising data protection and user privacy. The system has been designed considering two control strategies: 1) Edge Node Control (ENC), and 2) Edge node and Cluster Control (ECC). In both ENC and ECC, the energy supplier / distributor does not know about the specific control actions taking place within buildings. In ECC the distributor understands the impact of single buildings / agents in the network, while such knowledge is not needed in ENC.
[0097] In ENC, the need for data transfer between buildings (or agents) and the energy provider is at a minimum level, resulting in a minimum data transfer, maximum data security and maximum user privacy. Meanwhile, there is no need for a high computational power at the cluster level since only one flexibility signal is transferred to the whole grid per time step (e.g. every 15 minutes) without any decision making at the cluster level.
[0098] In ECC, the privacy measures at the building level are less strict than ENC and the computational power is stronger on the supply side. This creates the opportunity to optimize the distribution of signals for the buildings connected to the grid and transfer specific signals to different buildings. In ECC, the higher intelligence at the cluster level knows about the impact of signals on the energy performance of agents. For example, it knows applying the flexibility signal of 2 on building A will decrease its energy demand for around 20% for the coming time step, while signal 4 decreases for 50%. By having such knowledge, in ECC it is possible to reach an optimum distribution of signals over time for the whole stock (plurality of buildings or the considered grid). Below in the results section, only ENC is discussed and presented.
[0099] Integrating RL into decision-making in relation to the flexibility signal
[0100] The flexibility signal, or any other signal such as price signal, is or can be interpreted as a notion of the conditions in the energy grid. To effectively respond to this signal, the RL-based engine is designed for decision making that controls the behavior of each agent in the grid.
[0101] Integration of RL into decision making for energy optimization has been used previously.
[0102] However, a persistent challenge in the existing technologies is controlling a big number of agents together.
[0103] The novel elements that the present embodiments provide are:
[0104] 1) designing an RL-based approach that enables decentralized control of agents without data sharing and only using the flexibility signal as an input for agents; 2) removing the need for data sharing and storage at the multi-agent (or cluster) level; 3) updating the policy of agents automatically during certain periods to adapt faster to variations of the environment;
[0105] 4) considering upcoming conditions to prepare agents for the new environmental conditions;
[0106] 5) combining 3 and 4 to enhance the flexibility and adaptation agility of agents even faster. Element 1, consists of different features (see below) to enhance an ordinary RL engine into a functioning and light RL engine that controls the performance of each agent optimally without sharing any of its data outside its boundaries. These features result in having a scalable engine that can perform e.g. at the level of one building, multiple buildings, neighborhood etc. The features also enable multiple RL engines (or different agents) to collaborate towards reaching the optimal goal / performance at the cluster level (synergic performance of agents).
[0107] Element 1 - feature 1 - multidimensional policy matrix: RL works based on finding / making a policy that consists of relevant actions for the environmental conditions of the agent, helping the agent to reach its optimal performance. In ARLEM, the policy is made as a multidimensional matrix in a way that each dimension represents a key characteristic of the environment or agent. For example, in the tested ARLEM system, the policy matrix is 3D, containing three dimensions of: signal (e.g. flexibility signal), time (e.g. hour of the day) and set of actions (e.g. a set of actions that control temperature setpoint, ventilation rate and heating / cooling power).
[0108] Fig. l is a schematic representation of the multidimensional policy matrix in ARLEM, drawn for a 3D sample policy, containing dimensions of signal, time and action, which the latter includes set of actions (or a combination of control s / actions with different natures). At each time step, an action (i.e. set of different actions / controls) will be set depending on the received (environmental) signal (or flexibility signal). Since there can be multiple actions per time per signal (number of actions are defined according to the size of the library), the selected action depends on defining the selection criteria (which varies depending on the nature of the problem). However, all the actions stored per time per signal are those that fulfill the defined criteria (according to the reward function in the RL algorithm).
[0109] Element 1 - feature 2 - bundling actions over time: To speed up the optimization process and reaching an optimum policy faster, a feature is designed to increase the number of experiences per signal per time step (the arbitrary time step that is used for optimization purposes) which is based on gathering actions for similar conditions (time wise - can be changed based on the need, e.g. based on outdoor weather conditions) under one time step. For example, if the aim is increasing energy flexibility and climate resilience during very warm summers (or very cold winters), the whole season can be selected as the total time frame, resulting in 92 days for summer (June, July and August) and 90 days for winter (December, January and February). The whole season is divided into days (e.g. 92 days for summer) and each day into 24 hours (see Fig. 2). Assuming that the control time step is 15 minutes, under each hour there will be four actions (and flexibility signals or the state status). Therefore, the cumulative number of states for each separate hour (for example hour 15.00) will be 92 x 4 = 368. Such an approach helps to evaluate many actions over a short period and find the optimum ones, helping to reach an optimum policy faster. The selection of the time span is flexible and can be changed based on needs and conditions (e.g. grouping per month instead of season or based on the outdoor weather conditions).
[0110] Fig. 2 is a schematic diagram showing the status of the system during the considered time span being divided into days and specific hours of days which are bundled together. In the visualized case, each hour is divided into four 15-min periods, increasing the number of cases four times (i.e. each dot contains four sets of data).
[0111] Element 1 - feature 3 - universal signal generation: RL works based on reacting to the environmental conditions. In ARLEM, the environmental conditions are represented in one universal signal that is distributed throughout the whole network of agents (e.g. the flexibility signal or price signal). The process of generating the universal signal is flexible. For example, the original signal (e.g. price signal) can be used. Alternatively, the signal can be generated or translated to a digital signal, for example between 0-5. The digital signal may be defined in an algorithm to generate the suitable signal based on the needs and situation of the network. This works as a function in the algorithm that compares the current conditions (e.g. energy load or EL) and generates a relevant signal in a comparative process. The following is an example of a case when the global signal is a digit between 0-5.
[0112] function ELsig = EL_signal(EL, EL signals)
[0113] if EL < EL signals(l)
[0114] ELsig = 0;
[0115] elseif EL >= EL signals(l) & EL < EL_signals(2)
[0116] ELsig = 1;
[0117] elseif EL >= EL_signals(2) & EL < EL_signals(3)
[0118] ELsig = 2;
[0119] elseif EL >= EL_signals(3) & EL < EL_signals(4) ELsig = 3;
[0120] elseif EL >= EL_signals(4) & EL < EL_signals(5)
[0121] ELsig = 4;
[0122] elseif EL >= EL_signals(5)
[0123] ELsig = 5;
[0124] end
[0125] See also Fig. 3
[0126] E < EM can flexibility signal=0
[0127] E] [ean E < EA —> flexibility signal=l
[0128] EA < E < EB — > flexibility signal=2
[0129] EB < E < Ec — > flexibility signal=3
[0130] Ec < E < EMax —> flexibility signal=4
[0131] EMax < E OR 0.9*EMax < E — > flexibility signal=5
[0132] Element 2 - unique universal signal per time step:
[0133] Most of existing energy management systems (EMSs) are based on sending specific control signals per agent. This requires analyzing the performance of agents and consequently having access to their data.
[0134] Using Element 1, enables ARLEM to work only based on receiving one universal signal per time step, distributed through the whole network based on the universal state of the network. Agents in the network react to the signal and adapt an optimal control individually. As a result, there is no need for data sharing and storage from agents; they act only as the receivers of the universal signal. Each agent, depending on the signal and time, adopts its own optimal control strategy. This feature makes ARLEM an easy option for multiagent grids where the number of controllers and actuators grow exponentially.
[0135] Fig. 4 is a schematic representation of ARLEM.
[0136] Fig. 5 is a schematic illustration of ARLEM at the urban area, cluster node, and edge node (top) showing the Flexibility signal from DSO / TSO transferred to cluster nodes. Cluster node also transfers the signal to its edge nodes. Edge node communicates with the devices / controllers in the building to apply flexibility. The workflow and communications within the edge node is demonstrated inside a house (bottom-left) and the ARLEM algorithm is demonstrated in the flowchart (bottom-right) showing the control (output) and monitor (input) ports on the smart device as the edge node. Element 3 - Signal-Reward Conditional Selection:
[0137] In the common RL approaches, the rewarding actions (commonly those with the reward of one) are valid choices for the agent. This approach may postpone the synchronized adaptability of a multi-agent system and consequently decrease its agility to environmental changes, especially when variations of the environment are fast. For example, if buildings (which usually have high thermal mass) are considered as the agents in an energy system, the ordinary 0 and 1 rewarding system of RL can result in slow reaction of the system to weather shocks. A solution provided in ARLEM is based on defining a third reward, equal to 0.5, which relates to fulfilling a major goal (e.g. energy reduction) while not addressing another goal (e.g. indoor comfort). When the energy system is under high pressure (e.g. for signal 5, corresponding to high energy demand in the network), for certain agents (defined by the user), it is possible to select actions with reward 0.5, helping the system to pass the extreme condition faster without the risk of energy shortage. Such signal-reward conditional selection enables the system to act faster and become more resilient. This selection approach is flexible and can be designed based on the needs of the controlled system. For example, the rewarding system can be different (e.g. 0-0.3-0.5-0.8-1) in relation to the main goals / criteria in the controlled system (e.g. energy saving, indoor comfort, CO2 reduction etc.).
[0138] Element 4 - Smooth Action Selection:
[0139] Agents in a multi-agent system can have multiple sets of actions that vary between agents. For example, one building may have only two control actions (e.g. setting heating setpoint and ventilation rate), while another building may have six (adding e.g. cooling setpoint, shading, window opening and other appliances). For many controllers, it is needed to avoid sudden jumps in the system and have smooth transitions. This is enabled in ARLEM by a tagging system in the algorithm, limiting / allowing jumps between actions. For example, when the limit is set to 2, the controller can pick only the two nearest actions after and before the current condition (e.g. if the heating setpoint is 19°C and assuming a one-degree jump, the next setpoint can be either 17°C, 18°C, 20°C or 21 °C). Thanks to the tagging system, such selection criteria can be different for different controllers (and agents), enhancing the freedom of setting systems and selection in ARLEM.
[0140] Element 5 - Adaptive RL:
[0141] RL is based on creating a policy, including set of optimal actions for agents, to perform optimally in an environment. In common RL approaches, reaching an optimum policy can take some time while a central controller is needed. In ARLEM, finding the optimum policy is fast thanks to Element 1, while no central controller (or a decision maker at a higher level) is needed. Since the environmental conditions change over time (e.g. weather conditions change at different temporal scales), agents may encounter unfamiliar situations (unprecedented conditions), leading to a lack of preparedness and suboptimal performance. Therefore, the policy may need to be updated after a certain period of time or whenever needed. A feature that is designed in ARLEM is updating the policy automatically after a certain period (e.g. after 15 days) or when a certain condition occurs (e.g. reduced energy saving or sudden shocks such as extreme climate variations). This feature is designed in a way to act faster than building a policy from scratch by removing some of the weak actions (e.g. the 10 weakest actions out of 24 actions) and replacing them with updated actions that fit better the new conditions, resulting in an updated policy. The interval and number of the updates are flexible and can differ between agents in the same grid (e.g. removing 5 actions every 10 days for agent 1 and 10 actions every month for agent 2). This occurs for each signal and time slot separately and independently. For example, if signals 4 and 5 happen more often during time slots 2-6, action libraries for these cases will become updated more often (if needed).
[0142] As shown in Fig. 6, the adaptive RL engine enables ARLEM to update the policy over time by replacing the weaker action sets with those better suited for the new environmental conditions. This occurs for each signal and time slot separately and independently, meaning that the adaptation / update happens when is needed and is not forced all the time and for all the conditions.
[0143] Element 6 - Predictive RL: This element enables ARLEM (called PARLEM when the predictive feature is also being added) to prepare for the upcoming conditions (e.g. the next 24 hours). This works by considering the upcoming conditions (e.g. the weather forecast for the next 24 hours) in the signal generation function. In this regard, another component is added to the function for generating the global signal, reflecting upon the upcoming conditions. The following example shows a case to cope with extreme cold conditions beforehand (24 hours ahead). As visible, the definition of the function is flexible and can be adjusted for the needs of the system.
[0144] function ELsig = EL_signal_Tout(EL, EL signals, Tout)
[0145] if EL < EL signals(l)
[0146] ELsig = 0;
[0147] el seif CL >= EL signals(l) & EL < EL_signals(2)
[0148] ELsig = 1;
[0149] el seif EL >= EL_signals(2) & EL < EL_signals(3)
[0150] ELsig = 2; elseif EL >= EL_signals(3) & EL < EL_signals(4)
[0151] ELsig = 3;
[0152] elseif EL >= EL_signals(4) & EL < EL_signals(5)
[0153] ELsig = 4;
[0154] elseif EL >= EL_signals(5)
[0155] ELsig = 5;
[0156] end
[0157] ifTout<-12
[0158] ELsig = ELsig + 2;
[0159] elseif Tout<-8
[0160] ELsig = ELsig + 1;
[0161] end
[0162] if ELsig > 5
[0163] ELsig = 5;
[0164] end
[0165] Element 7 - Predictive Adaptive RL: This element is the combination of elements 6 and 7, enabling PARLEM to updates its policy while being informed about the upcoming conditions. Due to the high flexibility of elements 6 and 7, the flexibility of element 7 is also high and proper settings can be set depending on the need and performance.
[0166] Results
[0167] Results are presented for two case studies that ARLEM is used to control the energy performance of typical neighborhoods in Madrid and Stockholm, each having 24 buildings that include four different building types. In all the following cases, ARLEM policy is constructed by gathering 24 sets of actions per signal per hour. For the adaptive ARLEM and predictive PARLEM, results are only presented for Stockholm.
[0168] Summer in Madrid
[0169] In Madrid, the performance of ARLEM is assessed during a typical summer and a summer with two heatwaves (called ‘hot summer’ hereafter). The first heatwave starts on the 52nd day in summer and lasts for about six days, while the second one starts on the 73rd day and lasts for about 11 days.
[0170] The performance of ARLEM without adaptive and predictive features is assessed during the hot summer in comparison to the ordinary energy control system during typical and hot summers using boxplots and percentiles over three months in Fig. 7. Among the three summer months, June represents only typical conditions since the first heatwave happens at the end of July and the second and long one in the middle of August. Although there is no heatwave in June, ARLEM decreases the cooling demand both on the average and peak levels. The impact of ARLEM is very visible in August, which accommodates the longest heatwave. ARLEM decreases both the average (for about 8%) and peak (for about 20-25%) cooling demands, which for the latter the peak loads decrease to values very close to typical summer conditions, effectively contributing into peak shaving.
[0171] In Fig. 7, there is shown energy demand during summer with (HW) and without (TDY) heatwave in Madrid when the energy control is performed with ARLEM and without (Ref). In this case, the ARLEM control does not update its policy and there is no control limit.
[0172] Fig. 8 shows cumulative cooling demand and degree days over summer in Madrid for different modes of running ARLEM. ‘No PU’ means there is no policy update while ‘PU’ means there is policy update by updating n policies (e.g. lOp or 10 policies) every x day (e.g. 30d or 30 days). Three control criteria are compared, which are ‘Jump’ (no criteria and complete freedom to select any control action), ‘TwoStep’ (selecting among the two actions before and after the current action), and ‘OneStep’ (selecting only among the next and previous actions).
[0173] Performance of ARLEM with and without adaptive RL are compared for some specific cases in Fig. 8 by checking the cumulative cooling demand and degree days, which the latter is the multiplication of the number of days and discomfort temperatures (number of degrees out of the comfort range). In ‘No PU’ cases, the adaptive RL is not activated while in ‘PU’ the RL engine updates its policy by replacing p policies (e.g. lOp or 10 policies) every d days (e.g. 30d or 30 days). Three control criteria are compared, which are ‘Jump’ (no criteria and complete freedom to select any control action), ‘TwoStep’ (selecting among the two nearest actions before and after the current action), and ‘OneStep’ (selecting only among the next and previous actions). The heatwave periods are marked by rectangles in all the figures.
[0174] For the Jump case in Fig. 8, ARLEM decreases the cooling demand for about 9% while the difference is small between PU and No PU cases, both for energy demand and indoor comfort. In this case, updating policy when there are no control criteria does not change the performance of ARLEM considerably. For the TwoStep case, No PU decrease the cooling demand for about 9% while updating the policy results in saving energy for about 20%. Such an increase in energy saving comes with lowering the comfort for about 82% in comparison to No PU (increasing from about 66 degree days to 110). In the case of OneStep, No PU decreases the cooling demand for about 9% while updating the policy saves 8% more energy (reaching 17% energy saving), increasing the degree days for about 44% in comparison to No PU. The performance of multiple ARLEM control strategies is compared in Fig. 9 by plotting the boxplots and calculating some statistics. The first two boxplots on the left show the energy demand for the reference building (Ref - without ARLEM) during typical (TDY) and heatwave (HW), followed by those controlled by different versions of ARLEM during summer with heatwave. According to the calculated percentiles in Fig. 9, all the ARLEM versions decrease the average energy demand and percentiles (from 85% and above) even to values lower than ‘TDY -Ref . Updating the policy over time and limiting the selection criteria helps in reducing the average and peak energy demand further. For example, comparing ‘No policy update’ and ‘ lOp 30d’ shows that combining policy update and selection criteria can result in reducing average and median to values even lower than the typical summer. The performance of ‘one step’ and ‘two steps’ selection criteria are not very different; the latter decreases the energy demand average and peaks a bit more. The difference is not also considerable between lOp and 20p update strategies as well as the updating period (30d, 15d or 7d). The major impact is from updating the policy and limiting the selection criteria together. The combination of two helps to boost the performance of ARLEM in saving energy and shaving peaks
[0175] Fig. 9 shows energy demand during summer with (HW) and without (TDY) heatwave for different policy update and control limit strategies of ARLEM. Black dots are average values and green are medians.
[0176] Fig. 10 shows distribution of the indoor temperature with ARLEM with three control limits and without any policy update.
[0177] Fig. 1 Ishowns distribution of the indoor temperature with ARLEM with 10P policy update and three control limits.
[0178] Fig. 12 shows distribution of the indoor temperature with ARLEM with 20P policy update and three control limits.
[0179] Impacts of different control strategies on the indoor comfort are studied in Fig. 10, Fig. 11, and Fig. 12 by comparing the distribution of indoor temperature over summer. Considering the range of 19°C-25°C as the comfort range in Madrid, around 64% of time the indoor temperature is kept in the comfort limit when policy is not being updated according to Fig. 10. Updating the policy with 10P and 20P strategies, plotted respectively in Fig. 11 and Fig. 12, changes the distribution; the overheating hours may increase up to 20% depending on adopted control strategy. For example, in Fig. 11 setting ‘no limit’ selecting criteria keeps the discomfort hours at the same level as not updating the strategy in Fig. 10, while limiting the selection to one or two steps can increase warm hours for about 20%. The one step selection criterion induces around 15% less hours above 27°C compared to the two step criterion. Winter in Stockholm
[0180] In Stockholm, the performance of ARLEM is assessed during a typical winter and a winter with two cold waves (called ‘cold winter’ hereafter). The first cold wave starts on the 45th day of winter and lasts for about six days, while the second one starts on the 72nd winter day and lasts for about 10 days.
[0181] Fig. 13 shows energy demand during winter with (CW) and without (TDY) cold wave in Stockholm when the energy control is performed with ARLEM and without (Ref). In this case, the ARLEM control does not update its policy and there is no control limit.
[0182] The performance of ARLEM without adaptive and predictive features is assessed during the cold winter in comparison to the ordinary energy control system during typical and hot summers using boxplots and percentiles over three months in Fig. 13. Among the three winter months, there is no cold wave in December, the first happens in January and the second in February. ARLEM decreases the heating demand both on the average and peak levels in December, improving the energy performance of buildings during also typical conditions. In January, which has the first and shorter cold wave, ARLEM decreases the average heating demand for 15% and the peak demands for 11-14%. These values for February are about 16% and 14-16%, respectively. ARLEM effectively contributes to both decreasing the average heating demand and peak shaving.
[0183] Fig. 14 shows cumulative heating demand (left) and degree days (right) over winter in Stockholm for different modes of running ARLEM. ‘No PU’ means there is no policy update while ‘PU’ means there is policy update by updating n policies (e.g. 1 Op or 10 policies) every x day (e.g. 30d or 30 days). Three control criteria are compared, which are ‘Jump’ (no criteria and complete freedom to select any control action), ‘TwoStep’ (selecting among the two actions before and after the current action), and ‘OneStep’ (selecting only among the next and previous actions).
[0184] The performance of ARLEM with and without adaptive RL are compared for some specific cases in Fig. 14 by plotting the cumulative heating demand and degree days, which the latter is the multiplication of the number of days and discomfort temperatures (the comfort range is considered 19-24°C in Stockholm). The cold wave periods are marked by rectangles in all the figures.
[0185] For the Jump case in Fig. 14, ARLEM decreases the heating demand for about 16% while the difference is small between PU and No PU cases. In this case, updating the policy every 30 days by replacing 20 sets of actions results in less degree days, keeping the indoor comfort at the defined level 15% more than the other options. For the TwoStep case, ARLEM saves energy for about 16%, while updating the policy (PU lOp 7d) can result in 19% more indoor comfort. For the OneStep case, ARLEM saves around 13% energy compared to ordinary control, while updating the policy (PU 20p 30d) results in 27% more indoor comfort. Obviously, ARLEM in general can reduce the heating demand during the cold winter to the levels of the typical winter, while updating the policy through implementing the adaptative RL results in better indoor comfort.
[0186] Fig. 15 shows energy demand during winter with (CW) and without (TDY) cold wave for different policy update and control limit strategies of ARLEM.
[0187] The performance of multiple ARLEM control strategies is compared in Fig. 15 by plotting the boxplots and showing some percentiles. The first two boxplots from the left show the energy demand for the reference building during typical (TDY) and cold winter (CW), followed by those controlled by different versions of ARLEM during cold winter. All the ARLEM versions decrease the average energy demand for about 15%, to levels very similar to the typical winter. They also successfully decrease the peak energy demand values, 10-15% lower than the values with the ordinary control. Among the considered cases, PU 20p 30d with TwoStep selection criterion shows the best performance.
[0188] Fig. 16 shows distribution of the indoor temperature with ARLEM with three control limits and without any policy update.
[0189] Fig. 17 shows distribution of the indoor temperature with ARLEM with 10P policy update and three control limits.
[0190] Fig. 18 shows distribution of the indoor temperature with ARLEM with 20P policy update and three control limits.
[0191] The distribution of the indoor temperature in the 24 studied buildings are plotted in Fig. 16, Fig. 17, and Fig. 18 respectively for ARLEM with no policy update, being updated by replacing 10 sets of actions (lOp) and 20 sets (20p). Considering that some of the buildings are very large and most of them have limited control options (e.g. limited maximum power and large thermal masses), having the minimum temperature of 17°C during very cold days is acceptable. As figures show, for all the ARLEM version, the indoor temperature stays in the range of 17-21 °C while the ‘One Step’ selection criterion provides the best comfort, as is visible for the 20p case in Fig. 18. Depending on the control potentials in buildings and user preferences, a suitable version of ARLEM can be selected to fulfill the energy and comfort goals.
[0192] The predictive features of ARLEM are investigated in the following, by running ARLEM for the cold winter in Stockholm with no weather forecast and using weather forecasts for the next 3, 6, 12, 24 and 48 hours. When controlling buildings (or any other mechanism), it is mostly desired to avoid frequent jumping of the controllers between control options. Focusing on the indoor setpoint temperature, Fig. 18 shows how setting the predictive feature of ARLEM can reduce the jumping of controllers. It is obvious by focusing on the second cold wave period, which the black line is flattened in the predictive cases compared to the ’No forecast’ case. In some cases, such as ‘24h forecast’, the variations are very much reduced. This is also obvious in Fig. 20 which shows the indoor temperature itself. Boxplots of the heating demand during cold wave periods in Fig. 22 show that the average energy demand decreases by having the predictive P ARLEM together with increasing the range of values (larger IQR).
[0193] The distribution of the indoor temperature during cold wave periods in Fig. 23 show that the predictive feature keeps the indoor temperature at the lower bound around 25% more than ‘No forecast’ case. Apparently, this results in saving more energy during extreme conditions. Looking into the whole season in Fig. 24, shows that the major impact is during extreme weather conditions (as was aimed and expected); overall the temperature distribution is acceptable in all the cases. A very interesting feature of ARLEM is that it learns about the thermal capacity of buildings without knowing anything about them (it is a model-free RL), which is being nicely used when the predictive function is enabled.
[0194] Fig. 19 shows temporal distribution of indoor setpoint temperatures when ARLEM runs without (no forecast) and with predictive function for five different forecast spans. The solid black lines represent the average values of all the buildings per time step and light grey lines the distribution of values per building.
[0195] Fig. 20 shows temporal distribution of indoor temperatures when ARLEM runs without (no forecast) and with predictive function for five different forecast spans. The solid black lines represent the average values of all the buildings per time step and light grey lines the distribution of values per building.
[0196] Fig. 21 shows temporal distribution of energy demands of buildings when ARLEM runs without (no forecast) and with predictive function for five different forecast spans. The solid black lines represent the average values of all the buildings per time step and light grey lines the distribution of values per building.
[0197] Fig. 22 shows distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0198] Fig. 23 shows distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM runs without (no forecast) and with predictive function for five different forecast spans. Fig. 24 shows distribution of the indoor temperature in all the buildings during winter when ARLEM runs without (no forecast) and with predictive function for five different forecast spans.
[0199] Fig. 25 shows distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with no selection criteria (Jump) runs with predictive function for four different forecast spans.
[0200] Fig. 26 shows distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with TwoStep selection criteria runs with predictive function for four different forecast spans.
[0201] Fig. 27 shows distribution of the hourly heating demand of the neighborhood during cold wave periods when ARLEM with OneStep selection criteria runs with predictive function for four different forecast spans.
[0202] Fig. 28 shows distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with no selection criteria (Jump) runs with predictive function for four different forecast spans.
[0203] Fig. 29 shows distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with TwoStep selection criteria runs with predictive function for four different forecast spans.
[0204] Fig. 30 shows distribution of the indoor temperature during cold wave periods in all the buildings when ARLEM with OneStep selection criteria runs with predictive function for four different forecast spans.
[0205] In the following, it is investigated how ARLEM performs when the adaptative and predictive features of ARLEM are implemented together (also called PARLEM). The focus is on the two cold wave periods in Stockholm considering four forecast spans of 3, 6, 12 and 24 hours. The adaptive RL is updating its policy by replacing 10 sets of actions (lOp) every 30, 15, 7, 3, 2, and 1 days. Three selection criteria with no limit (Jump), two nearest (TwoStep) and one nearest (OneStep) are used. The distribution of hourly energy demands for these selection criteria are plotted respectively in Fig. 25, Fig. 26 and Fig. 27, followed by the distribution of indoor temperature in Fig. 28, Fig. 29 and Fig. 30. In comparison with Fig. 22 and Fig. 23, which show the performance of predictive RL without policy update, some improvements are visible in the ARLEM performance, mainly for OneStep selection criteria, especially for indoor temperature in Fig. 30.
Claims
CLAIMS1. A computer-implemented method for controlling energy performance of an agent among a plurality of agents within an energy network in which energy is distributed to the plurality of agents, the method being performed by an optimization and control unit associated with the agent, the method comprising:receiving a flexibility signal from a central server system external to and operating above the plurality of agents, the flexibility signal being indicative of the total energy load of the plurality of agents in relation to a predefined reference value and reflecting a current or upcoming environmental condition, the flexibility signal being a digital signal; and operating a model-free reinforcement learning, RL, engine within the optimization and control unit to select a control action based on the received flexibility signal, a current time, and a multidimensional policy matrix, wherein the multidimensional policy matrix comprises a dimension representing the flexibility signal, a dimension representing time, and a dimension representing a set of control actions.
2. The method according to claim 1, further comprising:assigning a reward to possible control actions in the set of control actions using a reward mechanism, the reward mechanism being a conditional selection system which in addition to a reward of 1 for a control action achieving a goal and a reward of 0 for other control actions comprises at least one weighted reward control action which relates to a control action fulfilling a primary goal while compromising a secondary goal; andconditionally selecting a control action that corresponds to a weighted reward control action when the flexibility signal indicates extreme conditions on the energy network, thereby controlling energy performance of the agent within the energy network by reducing the overall energy demand of the agent and contributing to a system-wide response to extreme conditions.
3. The method according to claim 1 or 2, further comprising updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions.
4. The method according to claim 3, wherein updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions is performed periodically.
5. The method according to claim 3 or 4, wherein updating the multidimensional policy matrix by replacing a subset of weakest control actions with new control actions is performed upon a condition occur.
6. The method according to any one of claims 3-5, where the weakest control actions are identified as those with the lowest assigned reward.
7. The method according to any one of claims 1-6, wherein conditionally selecting a control action comprises selecting the control actions only from a predefined number of nearest control actions relative to a current control action.
8. The method according to any one of claims 1-7, wherein the flexibility signal is a price signal.
9. The method according to any one of claims 1-7, wherein generating the flexibility signal at the central server system comprises:defining a range of values that indicate the need for flexibility;comparing the current energy load of the plurality of agents to a set of predefined reference values for a specific time or season; andassigning a digital signal from the defined range based on the comparison.
10. The method according to any one of claims 1-9, further comprising generating the flexibility signal at the central server system by:obtaining real-time data related to at least one environmental factor, such as solar irradiance, wind speed, or ambient temperature, from one or more sensors;analyzing the real-time data to determine a current state of the environmental condition; andgenerating the flexibility signal based on the current state, wherein the signal represents the immediate availability or constraint of a resource influenced by the environmental condition, such as the generation capacity of a renewable energy source.
11. The method according to any one of claims 1-10, further comprising:receiving a predictive environmental signal indicative of at least one future environmental factor, such as predicted solar irradiance, forecasted temperature, forecasted wind speed, or anticipated cloud cover, from an external source; andmodifying the selection of the control action based on the predictive environmental signal.
12. The method according to any one of claims 1-11, wherein the flexibility signal is a single universal signal transferred to all agents in the network per time step.
13. An optimization and control unit associated with an agent among a plurality of agents within an energy network in which energy is distributed to the plurality of agents, the optimization and control unit is configured to control energy performance of the associated agent, the optimization and control unit comprising processing circuitry for carrying out the method according to any one of claims 1-7.
14. A computer program comprising instructions which, when the program is executed by optimization and control unit having processing capabilities, cause the optimization and control unit to perform the method according to any one of claims 1-7.
15. A computer-readable storage medium having stored thereon a computer program according to claim 14.
16. A system for controlling energy performance of a plurality of agents in an energy network, the system comprising:a central server system configured to generate a flexibility signal, the flexibility signal being a digital signal indicative of the total energy load of the plurality of agents in relation to a predefined reference value and reflecting a current or upcoming environmental condition; and a plurality of optimization and control units, each being assigned to a separate one of the plurality of agents, wherein each optimization and control unit is configured to:receive the flexibility signal,operate independently without awareness of other optimization and control units, operate a model-free reinforcement learning, RL, engine to select a control action based on the received flexibility signal, a current time, and a multidimensional policy matrix, wherein the multidimensional policy matrix comprises a dimension representing the flexibility signal, a dimension representing time, and a dimension representing a set of control actions.
17. The system according to claim 16, wherein each optimization and control unit is further configured to:assign a reward to possible control actions in the set of control actions using a reward mechanism, the reward mechanism being a conditional selection system which in addition to a reward of 1 for a control action achieving a goal and a reward of 0 for other control actions comprises at least one weighted reward control action which relates to a control action fulfilling a primary goal while compromising a secondary goal; andconditionally select an action corresponding to a weighted reward control action when the flexibility signal indicates extreme conditions within the energy network, thereby controlling energy performance of the agent by reducing its overall energy demand of the agent and contributing to a system-wide response to extreme conditions.
18. The system according to claim 17, wherein each optimization and control unit is further configured to update the multidimensional policy matrix by replacing a subset of weakest actions with new actions.
19. The system according to any one of claims 16-18, wherein each optimization and control unit is further configured to conditionally select a control action by selecting the control action only from a predefined number of nearest control actions relative to a current control action.
20. The system according to any one of claims 16-19, wherein generating the flexibility signal at the central server system comprises:define a range of values that indicate the need for flexibility;compare the current energy load of the plurality of agents to a set of predefined reference values for a specific time or season; andassign a digital signal from the defined range based on the comparison.
21. The system according to any one of claims 16-20, wherein generating the flexibility signal at the central server system comprises:obtain real-time data related to at least one environmental factor, such as solar irradiance, wind speed, or ambient temperature, from one or more sensors;analyze the real-time data to determine a current state of the environmental condition; andgenerate the flexibility signal based on the current state, wherein the signal represents the immediate availability or constraint of a resource influenced by the environmental condition, such as the generation capacity of a renewable energy source.
22. The system according to any one of claims 16-21, wherein each optimization and control unit is further configured to:receive a predictive environmental signal indicative of at least one future environmental factor, such as predicted solar irradiance, forecasted temperature, forecasted wind speed, or anticipated cloud cover, from an external source; andmodify the selection of the control action based on the predictive environmental signal..
23. The system according to any one of claims 16-22, wherein the flexibility signal is a single universal signal transferred to all agents in the network per time step.
24. The system according to any one of claims 16-23, wherein each of the plurality of agents is a building, a part of a building or part of the energy network.
25. The system according to any one of claims 16-24, wherein a same flexibility signal is provided to all optimization and control units for providing demand-side management.