Methods and systems for training building control using simulated and real experience data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

HVAC systems face challenges in accurately training reinforcement learning models due to the need for substantial data, where simulated experience data may not capture actual operations, and real experience data collection is slow, requiring a method to improve both short-term and long-term control efficiency.

Innovation Solution

A method that combines simulated and real experience data to train reinforcement learning models, using a dynamic model to generate initial simulated data and subsequently retraining the model with real data, potentially retraining the dynamic model to generate additional simulated data for continuous improvement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If simulated experience data is used to train the reinforcement learning model, then training speed is improved, but data accuracy deteriorates

Engineering Contradiction:
Improvetraining speedVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent combines simulated experience data and real experience data into a unified training dataset. The reinforcement learning model is trained initially on simulated data for fast iteration, then retrained on real data collected from actual HVAC system operation to improve accuracy. This merging approach resolves the contradiction by leveraging the speed advantage of simulation and the accuracy advantage of real data.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary training using simulated experience data before deploying the model to real HVAC systems. This preliminary action allows the model to learn basic control policies quickly without waiting for slow real-data collection, then refines the policies using actual operational data afterward.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If real experience data is collected to train the reinforcement learning model, then data accuracy is improved, but training time increases

Engineering Contradiction:
Improvedata accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges simulated experience data (fast to generate) with real experience data (accurate but slow to collect) into a combined training dataset. This allows the system to achieve high data accuracy while minimizing training time by not relying solely on slow real-data collection.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system continuously collects real experience data from HVAC operation and continuously retrains the model, creating an ongoing improvement cycle. This continuous process ensures the model benefits from accumulating real data over time without requiring lengthy batch training periods.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If the reinforcement learning model is retrained frequently with real data, then control policy accuracy is improved, but computational resources are consumed

Engineering Contradiction:
Improvecontrol policy accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system implements periodic retraining of the reinforcement learning model using real experience data collected during HVAC operation. Instead of continuous retraining, the model is retrained at intervals when sufficient real data has been accumulated, balancing accuracy improvement with computational resource conservation.

Inventive Principle:
Principle #19Periodic action

4Quantity of substance

If simulated experience data is used, then data availability is improved, but realism deteriorates

Engineering Contradiction:
Improvedata availabilityVSAvoidrealism
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent combines simulated experience data (abundant but less realistic) with real experience data (limited but highly realistic) to create a comprehensive training dataset. This merging ensures both data availability and realism are achieved.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses simulation models to create copies of HVAC system behavior for generating training data. These simulated copies provide abundant training examples, which are then refined by incorporating actual sensor data from real systems to improve realism.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11859847B2Methods and systems for training building control using simulated and real experience data
Publication Date: 2024.01.02 TYCO FIRE & SECURITY GMBH
  • US11859847B2 patent drawing
  • US11859847B2 patent drawing
  • US11859847B2 patent drawing

AI summary

Systems and methods for training a reinforcement learning (RL) model for HVAC control are disclosed herein. Simulated experience data for the HVAC system is generated or received. The simulated experience data is used to initially train the RL model for HVAC control. The HVAC system operates within a building using the RL model and generates real experience data. A determination may be made to retrain the RL model. The real experience data is used to retrain the RL model. In some embodiments, both the simulated and real experience data are used to retrain the RL model. Experience data may be sampled according to various sampling functions. The RL model may be retrained multiple times over time. The RL model may be retrained less frequently over time as more real experience data is used to train the RL model.