Reinforcement Learning Control for High Dimensional Systems Modeled by Partial Differential Equations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional air-conditioning systems face challenges in real-time control due to the complexity and variability of airflow models, which are often of infinite dimension and influenced by uncertain physical parameters, leading to inefficiencies and suboptimal energy consumption.

Innovation Solution

A system and method that utilize a reduced order model based on Boussinesq equations, incorporating a robust closure model to handle uncertainties, combined with reinforcement learning for real-time control, to optimize airflow management in HVAC systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a physical model of airflow is used to control the air-conditioning system, then the heat load requirements can be met, but the model is of infinite dimension and too complex for real-time control applications

Engineering Contradiction:
Improveheat load requirement satisfactionVSAvoidmodel dimension and complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the infinite-dimensional airflow model into a finite-dimensional reduced order model by identifying and retaining only the most influential parameters and modes that capture the essential airflow behavior. This segmentation allows real-time control while maintaining sufficient accuracy for heat load satisfaction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the critical dynamic characteristics from the complex physical model to create a simplified reduced order model. By taking out only the essential parameters needed for control decisions, the system achieves real-time computational feasibility without completely discarding the physics-based foundation.

Inventive Principle:
Principle #2Taking out (Extraction)

2Device complexity

If a simplified reduced order model is used for real-time control, then the computational complexity is reduced, but the model may not accurately capture the infinite dimension physical airflow dynamics

Engineering Contradiction:
Improvecomputational complexityVSAvoidairflow dynamics accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transforms the model representation by changing parameters from a complete infinite-dimensional description to a finite set of reduced parameters that capture dominant airflow modes. This parameter reduction maintains accuracy for control-relevant quantities while enabling real-time computation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a reduced order model as an intermediary between the complex physical reality and the control system. This intermediary model preserves the essential dynamics needed for control while being computationally tractable, acting as a bridge that maintains accuracy without requiring full physical model complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of energy

If conventional model-based control methods are used, then energy efficiency can be optimized, but the models often ignore installation-specific characteristics such as room size, causing deviation from actual system operation

Engineering Contradiction:
Improveenergy consumptionVSAvoidinstallation-specific characteristics adaptation
Core Design Contradiction:
Loss of energyVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic reduced order model that adapts to installation-specific characteristics such as room size and geometry. The model parameters are configured based on actual installation data, allowing the control system to maintain energy efficiency while accounting for specific building characteristics that conventional generic models ignore.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250264237A1Reinforcement Learning Control for High Dimensional Systems Modeled by Partial Differential Equations
Publication Date: 2025.08.21 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US20250264237A1 patent drawing
  • US20250264237A1 patent drawing
  • US20250264237A1 patent drawing

AI summary

An optimization controller is provided for controlling an operation of a heating, ventilation and air conditioning (HVAC) system for air-conditioning a room. The contoroller receives setpoint values, and system measurements and airflow measurements respectively from system sensors arranged in the HVAC system and airflow sensors arranged in the room and performs, by using a memory and a processor, determining a horizon value N based on the setpoints and n parameter values the hight dimensional physics-based model based on the system measurements, computing state trajectories corresponding to the parameter values, providing a set of RL controllers represented by first and second gains, the state trajectories, and the reference signals, performing warm-start of the RL policy gradient algorithm using the RL controllers with an initial gain, computing feedback gains, for the set of RL controllers according to the RL policy gradient algorithm, determining optimal feedback gains by averaging the feedback gains,generating a control command based on the optimal feedback gains, and control the operation of the of the HVAC system based on the generated command.