Reinforcement Learning Reduced-Order Estimator for HVAC Flow Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing control methods for dynamical systems, especially nonlinear systems like HVAC units, face challenges in accurately capturing the physical dynamics, leading to suboptimal control policies and instability.

Innovation Solution

The use of a reduced order model combined with a virtual control term, known as a closure model, is proposed. This model is updated using reinforcement learning to mimic the pattern of dynamics, allowing for more efficient and stable control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a reduced order model combined with a closure model is used, then computational complexity is reduced and control design is simplified, but accuracy in capturing physical dynamics may be compromised

Engineering Contradiction:
Improvecomputational complexityVSAvoidaccuracy of physical dynamics
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The closure model acts as an intermediary component that bridges the reduced order model and the actual system dynamics. It captures the discrepancy between the simplified model predictions and the true physical behavior, allowing the complex dynamics to be approximated through a combination of the computationally efficient reduced order model and the corrective closure model, thus resolving the contradiction between computational simplicity and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The invention transforms the problem from directly modeling complex physical dynamics with full-order models to using parameter-based closure models that adapt to capture dynamic behavior. By changing the approach from direct physical modeling to parameter-driven correction terms, the system achieves both computational efficiency and accuracy in representing system dynamics.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If data-driven control methods are used without physical models, then control policies can be determined from operational data, but the resulting black box controller does not consider system physics and cannot be influenced by control designers

Engineering Contradiction:
Improveease of control policy designVSAvoidability to incorporate physical dynamics
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The invention merges data-driven control approaches with physics-based modeling by combining the closure model (learned from operational data) with the reduced order physical model. This hybrid approach allows control designers to incorporate both empirical observations and physical principles, creating controllers that are both data-adapted and physically meaningful, thus resolving the contradiction between ease of data-driven design and ability to incorporate system physics.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The closure model provides a feedback mechanism that continuously corrects the reduced order model predictions based on the discrepancy between model outputs and actual system behavior. This feedback loop enables the controller to adapt to data-driven patterns while maintaining physical consistency, allowing control designers to influence the control policy through the structured closure model formulation.

Inventive Principle:
Principle #23Feedback

3Reliability

If indirect data-driven control methods are used to construct system models, then model-based control design is enabled, but large quantities of data are required and estimated models do not capture system physics

Engineering Contradiction:
Improvereliability of model-based controlVSAvoidquantity of operational data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of attempting to build a complete system model from scratch using large quantities of data, the invention uses partial modeling through the reduced order model and supplements it with a closure model that captures the essential dynamics. This partial action approach enables model-based control with reliable performance while requiring significantly less operational data compared to full data-driven model construction.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12313276B2Time-varying reinforcement learning for robust adaptive estimator design with application to HVAC flow control
Publication Date: 2025.05.27 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US12313276B2 patent drawing
  • US12313276B2 patent drawing
  • US12313276B2 patent drawing

AI summary

A computer-implemented method using a reinforcement learning trained reduced order estimator (RL-trained ROE) and a closure model is provided for controlling a heating, ventilation, and air conditioning (HVAC) system including actuators. The method uses a processor coupled with a memory storing instructions implementing the method, wherein the instructions, when executed by the processor, carry out at steps of the method, includes acquiring setpoints of the HVAC system from a user input and measurement data from sensors arranged in the HVAC system,computing a high-dimensional state estimate using the measurement data and an estimate of reduced-order state from the RL-trained ROE, determining a controller with respect to the setpoints by using the RL-trained ROE, generating control commands based on the controller, and transmitting the control commands to the actuators of HVAC system via an output interface.