Air Conditioning Control Using Delayed-Reward Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Air conditioning devices face challenges in optimal control due to delayed responses between control actions and output changes, making it difficult to implement effective control using conventional reinforcement learning methods.

Innovation Solution

A method using a reinforcement learning agent to determine control actions and rewards for air conditioning devices, accounting for reward delay times and action maintenance times to optimize control actions, allowing for stable and efficient control by excluding irrelevant time points.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional reinforcement learning is used to control air conditioning devices, then control automation is achieved, but control accuracy deteriorates due to delayed response between control actions and output changes

Engineering Contradiction:
Improvecontrol automationVSAvoidcontrol accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by predicting future output values based on current control actions and historical data. The reinforcement learning model performs preliminary calculations to anticipate the delayed response effects, allowing the system to pre-adjust control actions to achieve more accurate temperature control despite the inherent time delay in the air conditioning system's response.

Inventive Principle:
Principle #10Preliminary action

2Use of energy by moving object

If reinforcement learning is applied to air conditioning control, then energy efficiency can be improved, but control stability deteriorates due to reward delay time

Engineering Contradiction:
Improveenergy efficiencyVSAvoidcontrol stability
Core Design Contradiction:
Use of energy by moving objectVSStability of the object's composition

Solution Approach 1:

The patent implements feedback by continuously monitoring actual output values, comparing them with predicted values, and using the differences to update the reinforcement learning model. This closed-loop feedback mechanism compensates for the reward delay time issue, maintaining control stability while enabling energy efficiency improvements through learned optimal control policies that adapt to the system's delayed response characteristics.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The model performs preliminary prediction of system behavior to anticipate delayed responses, allowing energy-efficient control actions to be planned in advance while maintaining stability through proactive adjustment rather than reactive correction after delays occur.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If control actions are adjusted frequently to account for delays, then control accuracy improves, but device complexity increases

Engineering Contradiction:
Improvecontrol accuracyVSAvoidcontrol system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses copying by creating a virtual model or simulation of the air conditioning system's behavior within the reinforcement learning framework. This digital twin or predictive model replicates the system's delayed response characteristics, allowing complex control adjustments to be calculated and optimized in the virtual model before being applied to the physical system, thereby achieving high control accuracy without proportionally increasing physical device complexity.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12188672B2Method for controlling air conditioning device based on delayed reward
Publication Date: 2025.01.07 MAKINAROCKS CO LTD
  • US12188672B2 patent drawing
  • US12188672B2 patent drawing
  • US12188672B2 patent drawing

AI summary

Disclosed is a method for controlling an air conditioning device, which is performed by at least one computing device, which includes: determining a control action for the air conditioning device at a first time point by using a reinforcement learning agent; determining a reward for the control action at the first time point based on a reward delay time by using the reinforcement learning agent; and performing reinforcement learning related to the control of the air conditioning device based on the determined reward, in which a time point when the reward delay time elapses from the first time point corresponds to a second time point, and the reward for the control action at the first time point is calculated while excluding situations after the first time point and before the second time point.