Reinforcement Learning Control for High Dimensional Systems Modeled by Partial Differential Equations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional air-conditioning systems face challenges in real-time control due to the complexity and variability of airflow models, which are often of infinite dimension and influenced by uncertain physical parameters, leading to inefficiencies and suboptimal energy consumption.
Innovation Solution
A system and method that utilize a reduced order model based on Boussinesq equations, incorporating a robust closure model to handle uncertainties, combined with reinforcement learning for real-time control, to optimize airflow management in HVAC systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a physical model of airflow is used to control the air-conditioning system, then the heat load requirements can be met, but the model is of infinite dimension and too complex for real-time control applications
Solution Approach 1:
The patent segments the infinite-dimensional airflow model into a finite-dimensional reduced order model by identifying and retaining only the most influential parameters and modes that capture the essential airflow behavior. This segmentation allows real-time control while maintaining sufficient accuracy for heat load satisfaction.
Solution Approach 2:
The patent extracts the critical dynamic characteristics from the complex physical model to create a simplified reduced order model. By taking out only the essential parameters needed for control decisions, the system achieves real-time computational feasibility without completely discarding the physics-based foundation.
2Device complexity
If a simplified reduced order model is used for real-time control, then the computational complexity is reduced, but the model may not accurately capture the infinite dimension physical airflow dynamics
Solution Approach 1:
The patent transforms the model representation by changing parameters from a complete infinite-dimensional description to a finite set of reduced parameters that capture dominant airflow modes. This parameter reduction maintains accuracy for control-relevant quantities while enabling real-time computation.
Solution Approach 2:
The patent introduces a reduced order model as an intermediary between the complex physical reality and the control system. This intermediary model preserves the essential dynamics needed for control while being computationally tractable, acting as a bridge that maintains accuracy without requiring full physical model complexity.
3Loss of energy
If conventional model-based control methods are used, then energy efficiency can be optimized, but the models often ignore installation-specific characteristics such as room size, causing deviation from actual system operation
Solution Approach 1:
The patent implements a dynamic reduced order model that adapts to installation-specific characteristics such as room size and geometry. The model parameters are configured based on actual installation data, allowing the control system to maintain energy efficiency while accounting for specific building characteristics that conventional generic models ignore.
Data Source
AI summary
An optimization controller is provided for controlling an operation of a heating, ventilation and air conditioning (HVAC) system for air-conditioning a room. The contoroller receives setpoint values, and system measurements and airflow measurements respectively from system sensors arranged in the HVAC system and airflow sensors arranged in the room and performs, by using a memory and a processor, determining a horizon value N based on the setpoints and n parameter values the hight dimensional physics-based model based on the system measurements, computing state trajectories corresponding to the parameter values, providing a set of RL controllers represented by first and second gains, the state trajectories, and the reference signals, performing warm-start of the RL policy gradient algorithm using the RL controllers with an initial gain, computing feedback gains, for the set of RL controllers according to the RL policy gradient algorithm, determining optimal feedback gains by averaging the feedback gains,generating a control command based on the optimal feedback gains, and control the operation of the of the HVAC system based on the generated command.


