Multi-agent deep reinforcement learning for dynamically controlling electrical equipment in buildings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for controlling building energy consumption, such as schedule-based and model predictive control, are sub-optimal due to restrictive deterministic models and computational complexity, especially in handling stochastic systems and continuous state spaces, and fail to account for real-world performance factors affecting entire building scenarios.

Innovation Solution

Implementing multi-agent deep reinforcement learning to dynamically control electrical equipment in buildings by training multiple RL agents using simulation models, where each agent monitors states affecting performance and learns optimal control parameters through a reward function combining energy components and penalties, allowing for efficient online learning and global optimal control parameter estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If model predictive control is used to optimize building energy consumption, then energy efficiency is improved, but the complexity of developing and calibrating accurate building models increases significantly

Engineering Contradiction:
Improvebuilding energy consumptionVSAvoidmodel calibration complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent replaces traditional model-based control mechanisms with a reinforcement learning-based control system. Instead of relying on complex calibrated building models for predictive control, the system uses RL agents that learn optimal control strategies through interaction with the building environment, substituting mechanical model calibration with adaptive learning algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The reinforcement learning system enables the building control system to self-optimize without requiring external model calibration expertise. The RL agents autonomously learn the building's thermal dynamics and optimal control strategies through trial and error, allowing the system to serve itself rather than requiring external modeling services.

Inventive Principle:
Principle #25Self-service

2Device complexity

If deterministic model assumption is used in MPC, then computational optimization is simplified, but the system becomes restrictive and cannot handle stochastic variations in building environments

Engineering Contradiction:
Improvecontrol algorithm complexityVSAvoidhandling of stochastic systems
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic adaptability to the control system by using reinforcement learning agents that continuously learn and adapt to changing building conditions. Instead of relying on static deterministic models, the RL system dynamically adjusts its control strategy based on real-time environmental variations, occupancy patterns, and building responses, enabling handling of stochastic systems.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If continuous state spaces are used in building control models, then measurement precision is improved, but traditional dynamic programming techniques become infeasible

Engineering Contradiction:
Improvetemperature and humidity measurement precisionVSAvoidcontroller design complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent substitutes traditional dynamic programming controllers with reinforcement learning-based controllers that are specifically designed to handle continuous state spaces. The RL approach uses function approximation techniques and neural networks to manage the complexity of continuous temperature, humidity, and other environmental parameters without requiring discretization of the state space.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Ease of operation

If single-agent reinforcement learning is used for building equipment control, then ease of operation is improved, but the system cannot account for multiple performance affecting factors across the entire building

Engineering Contradiction:
Improvecontrol system operation simplicityVSAvoidaccounting for multiple performance factors
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent segments the building control system into multiple independent reinforcement learning agents, each responsible for controlling specific equipment or zones. This segmentation allows each agent to focus on local optimization while the collective system accounts for building-wide performance factors. The multi-agent architecture maintains operational simplicity by keeping individual agents manageable while achieving comprehensive building optimization.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12111620B2Multi-agent deep reinforcement learning for dynamically controlling electrical equipment in buildings
Publication Date: 2024.10.08 TATA CONSULTANCY SERVICES LTD
  • US12111620B2 patent drawing
  • US12111620B2 patent drawing
  • US12111620B2 patent drawing

AI summary

Reinforcement Learning agent interacting with a real-world building to determine optimal policy may not be viable due to comfort constraints. Embodiments of the present disclosure provide multi-deep agent RL for dynamically controlling electrical equipment in buildings, wherein a simulation model is generated using design specification of (i) controllable electrical equipment (or subsystem) and (ii) building. Each RL agent is trained using simulation model and deployed in the subsystem. Reward function for each subsystem includes some portion of reward from other subsystem(s). Based on reward function of each RL agent, each RL agent learns an optimal control parameter during execution of RL agent in subsystem. Further, a global optimal control parameter list is generated using the optimal control parameter. The control parameters in the global optimal control parameters list are fine-tuned to improve subsystem's performance. Information on fine-tuning parameters of the subsystem and reward function are used for training RL agents.