Air Conditioning Control Using Delayed-Reward Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Air conditioning devices face challenges in optimal control due to delayed responses between control actions and output changes, making it difficult to implement effective control using conventional reinforcement learning methods.
Innovation Solution
A method using a reinforcement learning agent to determine control actions and rewards for air conditioning devices, accounting for reward delay times and action maintenance times to optimize control actions, allowing for stable and efficient control by excluding irrelevant time points.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional reinforcement learning is used to control air conditioning devices, then control automation is achieved, but control accuracy deteriorates due to delayed response between control actions and output changes
Solution Approach 1:
The patent applies preliminary action by predicting future output values based on current control actions and historical data. The reinforcement learning model performs preliminary calculations to anticipate the delayed response effects, allowing the system to pre-adjust control actions to achieve more accurate temperature control despite the inherent time delay in the air conditioning system's response.
2Use of energy by moving object
If reinforcement learning is applied to air conditioning control, then energy efficiency can be improved, but control stability deteriorates due to reward delay time
Solution Approach 1:
The patent implements feedback by continuously monitoring actual output values, comparing them with predicted values, and using the differences to update the reinforcement learning model. This closed-loop feedback mechanism compensates for the reward delay time issue, maintaining control stability while enabling energy efficiency improvements through learned optimal control policies that adapt to the system's delayed response characteristics.
Solution Approach 2:
The model performs preliminary prediction of system behavior to anticipate delayed responses, allowing energy-efficient control actions to be planned in advance while maintaining stability through proactive adjustment rather than reactive correction after delays occur.
3Measurement precision
If control actions are adjusted frequently to account for delays, then control accuracy improves, but device complexity increases
Solution Approach 1:
The patent uses copying by creating a virtual model or simulation of the air conditioning system's behavior within the reinforcement learning framework. This digital twin or predictive model replicates the system's delayed response characteristics, allowing complex control adjustments to be calculated and optimized in the virtual model before being applied to the physical system, thereby achieving high control accuracy without proportionally increasing physical device complexity.
Data Source
AI summary
Disclosed is a method for controlling an air conditioning device, which is performed by at least one computing device, which includes: determining a control action for the air conditioning device at a first time point by using a reinforcement learning agent; determining a reward for the control action at the first time point based on a reward delay time by using the reinforcement learning agent; and performing reinforcement learning related to the control of the air conditioning device based on the determined reward, in which a time point when the reward delay time elapses from the first time point corresponds to a second time point, and the reward for the control action at the first time point is calculated while excluding situations after the first time point and before the second time point.


