Methods and systems for training building control using simulated and real experience data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
HVAC systems face challenges in accurately training reinforcement learning models due to the need for substantial data, where simulated experience data may not capture actual operations, and real experience data collection is slow, requiring a method to improve both short-term and long-term control efficiency.
Innovation Solution
A method that combines simulated and real experience data to train reinforcement learning models, using a dynamic model to generate initial simulated data and subsequently retraining the model with real data, potentially retraining the dynamic model to generate additional simulated data for continuous improvement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If simulated experience data is used to train the reinforcement learning model, then training speed is improved, but data accuracy deteriorates
Solution Approach 1:
The patent combines simulated experience data and real experience data into a unified training dataset. The reinforcement learning model is trained initially on simulated data for fast iteration, then retrained on real data collected from actual HVAC system operation to improve accuracy. This merging approach resolves the contradiction by leveraging the speed advantage of simulation and the accuracy advantage of real data.
Solution Approach 2:
The system performs preliminary training using simulated experience data before deploying the model to real HVAC systems. This preliminary action allows the model to learn basic control policies quickly without waiting for slow real-data collection, then refines the policies using actual operational data afterward.
2Measurement precision
If real experience data is collected to train the reinforcement learning model, then data accuracy is improved, but training time increases
Solution Approach 1:
The patent merges simulated experience data (fast to generate) with real experience data (accurate but slow to collect) into a combined training dataset. This allows the system to achieve high data accuracy while minimizing training time by not relying solely on slow real-data collection.
Solution Approach 2:
The system continuously collects real experience data from HVAC operation and continuously retrains the model, creating an ongoing improvement cycle. This continuous process ensures the model benefits from accumulating real data over time without requiring lengthy batch training periods.
3Measurement precision
If the reinforcement learning model is retrained frequently with real data, then control policy accuracy is improved, but computational resources are consumed
Solution Approach 1:
The system implements periodic retraining of the reinforcement learning model using real experience data collected during HVAC operation. Instead of continuous retraining, the model is retrained at intervals when sufficient real data has been accumulated, balancing accuracy improvement with computational resource conservation.
4Quantity of substance
If simulated experience data is used, then data availability is improved, but realism deteriorates
Solution Approach 1:
The patent combines simulated experience data (abundant but less realistic) with real experience data (limited but highly realistic) to create a comprehensive training dataset. This merging ensures both data availability and realism are achieved.
Solution Approach 2:
The system uses simulation models to create copies of HVAC system behavior for generating training data. These simulated copies provide abundant training examples, which are then refined by incorporating actual sensor data from real systems to improve realism.
Data Source
AI summary
Systems and methods for training a reinforcement learning (RL) model for HVAC control are disclosed herein. Simulated experience data for the HVAC system is generated or received. The simulated experience data is used to initially train the RL model for HVAC control. The HVAC system operates within a building using the RL model and generates real experience data. A determination may be made to retrain the RL model. The real experience data is used to retrain the RL model. In some embodiments, both the simulated and real experience data are used to retrain the RL model. Experience data may be sampled according to various sampling functions. The RL model may be retrained multiple times over time. The RL model may be retrained less frequently over time as more real experience data is used to train the RL model.


