Learning device, air conditioning control system, inference device, air conditioning control device, and trained model generation method
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reinforcement learning for air conditioner control requires a long period to converge, leading to inadequate initial control performance.
Innovation Solution
A learning device that simulates thermal environments and employs reinforcement learning with simulated reward values to generate trained models for inferring air conditioner control values, reducing the learning period.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reinforcement learning is executed using actual measured data from the air conditioner and real environment, then the control accuracy and comfort level are improved, but the learning period becomes excessively long
Solution Approach 1:
The patent creates a virtual copy of the air conditioner system including the refrigeration cycle, indoor space thermal environment, and control mechanisms. This virtual model replicates the behavior of the actual system, allowing reinforcement learning to be performed on the copy rather than the real system. The virtual model includes components such as the compressor, expansion valve, heat exchangers, and thermal environment simulation, enabling realistic training without requiring actual hardware operation during the learning phase.
Solution Approach 2:
The patent performs reinforcement learning in advance using the virtual model before deploying the learned control policy to the actual air conditioner. The trained model is generated through preliminary training iterations in the virtual environment, accumulating learning results that can be directly applied to real-world operation. This preliminary action eliminates the need for lengthy on-site learning that would otherwise be required.
2Reliability
If reinforcement learning is performed in the real environment with actual air conditioner operation, then the learned policy reflects real-world conditions accurately, but the air conditioner cannot be appropriately controlled before convergence
Solution Approach 1:
The patent uses a comprehensive virtual model that accurately replicates the air conditioner's refrigeration cycle and thermal environment behavior. This copy allows the system to be controlled appropriately from the beginning using the trained model, while the virtual copy continues to serve as a training environment. The virtual model includes detailed representations of the compressor, expansion valve, heat exchangers, and indoor space thermal dynamics, ensuring the trained policy translates accurately to real-world conditions.
Solution Approach 2:
The patent introduces a trained model as an intermediary between the reinforcement learning process and the actual air conditioner control. This trained model, generated through virtual environment training, acts as a bridge that enables appropriate control of the real air conditioner from the start. The intermediary model translates the learned policies into actionable control commands for the actual system without requiring the real system to undergo the lengthy learning process.
3Measurement precision
If the reinforcement learning uses complex simulation of thermal environment, then the learning accuracy is improved, but the computational complexity and processing time increase
Solution Approach 1:
The patent employs simplified thermal environment simulation by changing key parameters from detailed physical measurements to representative variables. Instead of tracking every thermal parameter in the indoor space, the simulation uses essential parameters such as indoor temperature, outdoor temperature, and simplified thermal load representations. This parameter reduction maintains sufficient learning accuracy while dramatically reducing computational complexity and processing requirements.
Solution Approach 2:
The patent applies local quality by focusing simulation detail only where necessary for accurate learning. The virtual model incorporates detailed refrigeration cycle physics (compressor, expansion valve, heat exchanger dynamics) where precision is critical, while using simplified representations for less critical aspects such as overall indoor space thermal dynamics. This selective detail distribution maintains learning accuracy in critical areas without requiring excessive computational resources across the entire system.
Data Source
AI summary
A simulator of a learning device simulates a thermal environment of an indoor space predicted to result from air conditioning of the indoor space by an air conditioner in a situation in which at least one of a state of a refrigeration cycle included in the air conditioner and a state of the indoor space is given. A reinforcement learner executes reinforcement learning that employs, as a reward, a value based on the thermal environment simulated by the simulator, and thereby generates a trained model aimed at inferring, from the at least one of the state of the refrigeration cycle and the state of the indoor space, a control value of the air conditioner.


