System and method for training a control strategy for a heating, ventilation, and air-conditioning (HVAC) system

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning (RL) methods for training HVAC control strategies face challenges such as long training times, intermediate sub-optimal solutions, and limited scope in optimizing energy savings and user comfort, due to the use of synthetic data and reliance on other users' behavior without exploring all possible options.

Innovation Solution

A method involving federated learning to progressively refine HVAC control strategies from a common policy to individual user-specific policies, using user feedback on comfort and thermal models to determine rewards, and applying control actions in a hierarchical manner to minimize user discomfort and optimize performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning is used to train HVAC control strategies with synthetic data, then training speed is improved, but accuracy and reliability of the control strategy deteriorates

Engineering Contradiction:
Improvetraining speedVSAvoidaccuracy of control strategy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the reinforcement learning model using synthetic data generated from a surrogate model before deploying it with real HVAC units. This preliminary training phase allows the model to learn basic control patterns quickly without waiting for extensive real-world data collection, while subsequent fine-tuning with actual user feedback ensures the strategy adapts to real conditions and maintains accuracy.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If a single global control strategy is trained for all HVAC units, then device complexity is reduced, but adaptability to individual user needs deteriorates

Engineering Contradiction:
Improvecontrol strategy structureVSAvoiduser-specific comfort optimization
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies segmentation by dividing the population of HVAC users into distinct clusters based on their behavioral patterns, comfort preferences, and usage characteristics. Instead of applying a single uniform control strategy to all users, the system creates segmented user profiles and trains specialized control strategies for each segment, allowing the system to maintain relatively simple individual models while achieving high adaptability to diverse user needs through multi-group specialization.

Inventive Principle:
Principle #1Segmentation

3Reliability

If reinforcement learning explores all possible options to optimize energy savings and comfort, then solution quality is improved, but training time increases

Engineering Contradiction:
Improveoptimization qualityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by using a surrogate model to pre-generate synthetic training data that captures the essential dynamics of HVAC system behavior and user preferences. This allows the reinforcement learning algorithm to begin training with a comprehensive dataset that already encodes optimal or near-optimal control patterns, significantly reducing the exploration time needed while still achieving high-quality optimization results.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies copying by creating a surrogate model that replicates the complex HVAC system dynamics and user behavior patterns. This virtual copy allows the reinforcement learning algorithm to train extensively on synthesized data that mirrors real-world conditions without the time constraints of actual system operation, thereby achieving thorough optimization exploration in a compressed timeframe before deploying to real units.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4446824A1System and method for training a control strategy for a heating, ventilation, and air-conditioning (HVAC) system
Publication Date: 2024.10.16 ROBERT BOSCH GMBH
  • EP4446824A1 patent drawingFigure 1
  • EP4446824A1 patent drawingFigure 2
  • EP4446824A1 patent drawingFigure 3

AI summary

Aspects concern a method for training a control strategy for a heating, ventilation, and air-conditioning, HVAC, system, comprising applying control actions to a set of HVAC units comprising a plurality of HVAC units, gathering feedback from each HVAC unit of the set of HVAC units regarding results of the control actions, training a first HVAC control strategy using reinforcement learning, wherein rewards are at least partially determined based on the feedback and sub-dividing the set of HVAC units into different sub-sets of HVAC units and, for each sub-set, refining the first HVAC control strategy to a respective second HVAC control strategy for the sub-set using reinforcement learning, wherein rewards are at least partially determined based on the feedback gathered for HVAC units of the sub-set.