Air Conditioner Q-Learning Control for Adaptive Energy Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing air conditioner control methods lack adaptive adjustment capability, leading to high energy consumption, low control accuracy, and poor user experience due to manually set PID parameters, which fail to optimize operation under different conditions.
Innovation Solution
A control method using a reward matrix and Q-learning algorithm to optimize air conditioner operation by calculating the maximum expected benefit based on current state and action parameters, including compressor frequency, electronic expansion valve opening, and fan speed, with a constraint network model for default parameter filling and radial basis function neural network training for accuracy verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional PID control strategy with manually determined parameters is used, then the control method is simple to implement, but the adaptive adjustment capability is lacking and control accuracy is low
Solution Approach 1:
The patent applies dynamics by transitioning from static manually-determined PID parameters to dynamic adaptive parameters that automatically adjust according to real-time working conditions. The reinforcement learning agent continuously learns and updates control parameters based on environmental feedback, enabling the system to adapt to varying operating conditions and achieve high control accuracy without manual intervention.
Solution Approach 2:
The patent implements self-service through the reinforcement learning framework where the air conditioner system autonomously optimizes its own control parameters. The system independently collects data, trains the neural network model, and adjusts control strategies without external human input, achieving both high accuracy and automated operation.
2Use of energy by moving object
If traditional PID control with fixed parameters is used, then the system is easy to operate, but energy consumption is high and resource utilization is poor
Solution Approach 1:
The patent implements feedback mechanisms where the reinforcement learning agent continuously monitors system state and performance metrics, then adjusts control parameters accordingly. This closed-loop feedback enables the system to optimize energy consumption by adapting to real-time conditions, achieving both low energy usage and high automation level.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting control parameters based on learned patterns and real-time conditions. The neural network model continuously updates optimal parameter values, enabling the system to minimize energy consumption while maintaining effective automatic control across diverse operating scenarios.
3Adaptability or versatility
If manually determined PID parameters are used, then the control strategy is simple, but it cannot satisfy control requirements of different working conditions
Solution Approach 1:
The patent applies dynamics by transforming the control system from a static parameter configuration to a dynamic adaptive system. The reinforcement learning framework enables real-time parameter adjustment based on working conditions, achieving high adaptability while the modular architecture keeps implementation complexity manageable.
Solution Approach 2:
The patent implements preliminary action through offline training of the neural network model using historical data and simulations. This pre-learning phase prepares the system with prior knowledge, enabling it to quickly adapt to new working conditions during online operation, thus achieving high versatility without excessive real-time computational complexity.
4Adaptability or versatility
If reinforcement learning with Q-learning algorithm is implemented, then adaptive adjustment capability is improved, but the computational complexity increases
Solution Approach 1:
The patent applies preliminary action by performing extensive training computations offline before deployment. The Q-learning algorithm and neural network model are trained in advance using historical data and simulations, storing learned knowledge in the model parameters. During actual operation, the system only needs to infer from the pre-trained model, significantly reducing real-time computational complexity while maintaining high adaptability.
Solution Approach 2:
The patent replaces complex real-time computational mechanics with a pre-trained neural network inference system. Instead of performing heavy reinforcement learning computations during operation, the system uses the trained model to quickly determine optimal actions, substituting offline computational complexity for online simplicity while preserving adaptive capabilities.
Data Source
AI summary
The disclosure provides a control method and a device for an air conditioner. The method includes: a first reward matrix is constructed according to multiple sets of target operating parameters of an air conditioner, a maximum expected benefit of performing a current action in a current state is calculated based on the first reward matrix and a Q-learning algorithm, wherein the current state is represented by a current indoor environment temperature and a current outdoor environment temperature; target action parameters under the maximum expected benefit are acquired, and operation of the air conditioner is controlled based on second target action parameters, wherein the second target action parameters at least include a second target operating frequency of the compressor, a second target opening degree of the electronic expansion valve and a second target rotating speed of the external fan.


