Multi-agent deep reinforcement learning for dynamically controlling electrical equipment in buildings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for controlling building energy consumption, such as schedule-based and model predictive control, are sub-optimal due to restrictive deterministic models and computational complexity, especially in handling stochastic systems and continuous state spaces, and fail to account for real-world performance factors affecting entire building scenarios.
Innovation Solution
Implementing multi-agent deep reinforcement learning to dynamically control electrical equipment in buildings by training multiple RL agents using simulation models, where each agent monitors states affecting performance and learns optimal control parameters through a reward function combining energy components and penalties, allowing for efficient online learning and global optimal control parameter estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If model predictive control is used to optimize building energy consumption, then energy efficiency is improved, but the complexity of developing and calibrating accurate building models increases significantly
Solution Approach 1:
The patent replaces traditional model-based control mechanisms with a reinforcement learning-based control system. Instead of relying on complex calibrated building models for predictive control, the system uses RL agents that learn optimal control strategies through interaction with the building environment, substituting mechanical model calibration with adaptive learning algorithms.
Solution Approach 2:
The reinforcement learning system enables the building control system to self-optimize without requiring external model calibration expertise. The RL agents autonomously learn the building's thermal dynamics and optimal control strategies through trial and error, allowing the system to serve itself rather than requiring external modeling services.
2Device complexity
If deterministic model assumption is used in MPC, then computational optimization is simplified, but the system becomes restrictive and cannot handle stochastic variations in building environments
Solution Approach 1:
The patent introduces dynamic adaptability to the control system by using reinforcement learning agents that continuously learn and adapt to changing building conditions. Instead of relying on static deterministic models, the RL system dynamically adjusts its control strategy based on real-time environmental variations, occupancy patterns, and building responses, enabling handling of stochastic systems.
3Measurement precision
If continuous state spaces are used in building control models, then measurement precision is improved, but traditional dynamic programming techniques become infeasible
Solution Approach 1:
The patent substitutes traditional dynamic programming controllers with reinforcement learning-based controllers that are specifically designed to handle continuous state spaces. The RL approach uses function approximation techniques and neural networks to manage the complexity of continuous temperature, humidity, and other environmental parameters without requiring discretization of the state space.
4Ease of operation
If single-agent reinforcement learning is used for building equipment control, then ease of operation is improved, but the system cannot account for multiple performance affecting factors across the entire building
Solution Approach 1:
The patent segments the building control system into multiple independent reinforcement learning agents, each responsible for controlling specific equipment or zones. This segmentation allows each agent to focus on local optimization while the collective system accounts for building-wide performance factors. The multi-agent architecture maintains operational simplicity by keeping individual agents manageable while achieving comprehensive building optimization.
Data Source
AI summary
Reinforcement Learning agent interacting with a real-world building to determine optimal policy may not be viable due to comfort constraints. Embodiments of the present disclosure provide multi-deep agent RL for dynamically controlling electrical equipment in buildings, wherein a simulation model is generated using design specification of (i) controllable electrical equipment (or subsystem) and (ii) building. Each RL agent is trained using simulation model and deployed in the subsystem. Reward function for each subsystem includes some portion of reward from other subsystem(s). Based on reward function of each RL agent, each RL agent learns an optimal control parameter during execution of RL agent in subsystem. Further, a global optimal control parameter list is generated using the optimal control parameter. The control parameters in the global optimal control parameters list are fine-tuned to improve subsystem's performance. Information on fine-tuning parameters of the subsystem and reward function are used for training RL agents.


