Air conditioner control method and device based on model reinforcement learning, terminal and medium

By building an RC model and reinforcement learning agent, the Q table of the optimal air conditioning control strategy is trained, which solves the problems of high data demand and long training time in the traditional method, and achieves fast and efficient air conditioning demand response control, taking into account user comfort and energy efficiency.

CN120252126AActive Publication Date: 2025-07-04SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510744836.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Traditional reinforcement learning control methods have high data demand and long training time in building air conditioning systems, making it difficult to quickly and efficiently meet the practical application requirements of demand response control.

Method used

Using a model-based reinforcement learning method, a reinforcement learning agent is constructed by building an RC model that reflects the thermal behavior of the target building, and the agent is trained in the simulation environment generated by the RC model to obtain the Q table of the optimal air conditioning control strategy and deploy it to the air conditioning system for control.

Benefits of technology

Reliance on high-precision building physical models and massive historical data is reduced, and lightweight model construction and convenient deployment is realized. Efficient and accurate control strategies can be trained in a short time, taking into account user comfort and demand response goals, and achieving efficient, stable and highly robust air-conditioning demand response management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120252126A_ABST
    Figure CN120252126A_ABST
Patent Text Reader

Abstract

The invention provides an air conditioner control method and device based on model reinforcement learning, a terminal and a medium, and relates to the technical field of artificial intelligence. The method comprises the steps that an RC model reflecting the thermal behavior of a target building is built to serve as a training environment model; a reinforcement learning agent is constructed, the agent is trained in a reinforcement learning simulation environment generated based on an RC model, and a Q table with an optimal air conditioner control strategy is obtained; and the Q table is deployed into the air conditioning system carried on the target building, and operation of the air conditioning system is controlled through the intelligent agent based on the Q table with the optimal air conditioning control strategy. By combining the advantages of the RC model and reinforcement learning and utilizing information of reinforcement learning and environment interaction, the demand for historical data is reduced, an efficient and accurate control strategy can be trained in a short time, and the problems that a traditional reinforcement learning control method is high in data demand, long in training time and low in training efficiency are solved. And the actual application demand of demand response control is difficult to meet quickly and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to an air conditioner control method, device, terminal and medium based on model reinforcement learning. Background Art

[0002] Nowadays, low-carbon development has become a global consensus. The construction industry is an important part of global energy use and greenhouse gas emissions. Among them, the energy consumption of building air conditioning systems accounts for a large proportion, and there is huge energy-saving potential. However, with the increase in volatile new energy and the expansion of the peak-valley difference of the load, the stable and safe operation of the power grid has been impacted unprecedentedly. Therefore, as an important part of the total building energy consumption, the operation optimization control of building air conditioning systems is particularly important for reducing the energy consumption of the entire building air conditioning system and even the total building energy consumption. It has a high load regulation ability and is an important resource for realizing demand response. The demand response technology (DR, Demand Response) provides a solution for reducing building energy consumption and alleviating the pressure on the power grid.

[0003] However, many current traditional reinforcement learning control methods for building air conditioning systems have deficiencies such as high data requirements and long training times, and it is difficult to quickly and efficiently meet the actual application requirements of demand response control. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an air conditioner control method, device, terminal and medium based on model reinforcement learning for the above-mentioned defects of the prior art, aiming to solve the problems that traditional reinforcement learning control methods have high data requirements, long training times, and it is difficult to quickly and efficiently meet the actual application requirements of demand response control.

[0005] The technical solution adopted by the present invention to solve the technical problem is as follows: An air conditioner control method based on model reinforcement learning, wherein the method includes: Construct an RC model reflecting the thermal behavior of the target building as the training environment model; Construct a reinforcement learning agent, and train the reinforcement learning agent in a reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioner control strategy; Deploy the Q-table with the optimal air conditioner control strategy to the air conditioner system installed on the target building, and control the operation of the air conditioner system through the reinforcement learning agent based on the Q-table with the optimal air conditioner control strategy.

[0006] In one implementation, the RC model is composed of a thermal resistance and a heat capacity, and the expression of the RC model is: ; ; ; ; ; ; ; Among them, represents the heat transfer between the indoor air and the air conditioning system, represents the heat transfer between the building envelope and the indoor air, represents the heat transfer between the outdoor air and the building envelope, represents the indoor temperature, represents the air conditioning system temperature, represents the building envelope temperature, represents the outdoor temperature, represents the th moment of the air conditioning system temperature, represents the th moment of the air conditioning system temperature, represents the th moment of the indoor temperature, represents the th moment of the indoor temperature, represents the th moment of the building envelope temperature, represents the th moment of the building envelope temperature, represents the air conditioning system thermal resistance, represents the indoor air thermal resistance, represents the building envelope thermal resistance, represents the indoor heat capacity, represents the building heat capacity, represents the cooling power of the air conditioning system, represents the indoor heat, ONOFF represents the operating state of the compressor, t represents time in seconds, h represents hours, represents the mapping relationship between the indoor heat and hours h, represents the indoor heat corresponding to each hour.

[0007] In one implementation, the constructing the reinforcement learning agent includes: Determine the basic parameters of the agent and the environmental parameters, and construct the corresponding reinforcement learning agent according to the basic parameters of the agent and the environmental parameters; Among them, the basic parameters of the agent include the state space, action space, reward function, and ε-greedy policy setting value. The ε-greedy policy setting value is a constant between 0 and 1. The environmental parameters include outdoor temperature, indoor temperature, air conditioner operating status, electricity price, and time information; And, the Q-value update formula of the reinforcement learning agent is: ; Among them, s represents the state space, and the state space includes indoor temperature, outdoor temperature, and time information. represents the next state space after performing action a in state s. a represents the action space, and the action space is a binary discrete action set {0, 1}, where the element 0 represents the air conditioner is turned off, and the element 1 represents the air conditioner is turned on. represents the action space corresponding to the next state space r represents the reward function, and the reward function consists of temperature comfort penalty, electricity cost penalty, low temperature reward, high temperature reward, and precooling reward. represents the learning rate. represents the discount factor.

[0008] In one implementation, training the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioner control strategy includes: Collecting the historical outdoor temperature data and historical electricity price data of the target building; Using the historical outdoor temperature data and the historical electricity price data to drive the reinforcement learning agent to interact with the reinforcement learning simulation environment generated based on the RC model, so as to train the reinforcement learning agent to select actions based on the states obtained from the RC model, receive the rewards feedback by the environment, and update the Q-table until the training converges to obtain a Q-table with an optimal air conditioner control strategy.

[0009] In one implementation, before deploying the Q-table with the optimal air conditioner control strategy to the air conditioner system installed on the target building and controlling the operation of the air conditioner system by the reinforcement learning agent based on the Q-table with the optimal air conditioner control strategy, it further includes: Deploying the Q-table with the optimal air conditioner control strategy to the reinforcement learning simulation environment generated based on the RC model that is consistent with the training environment parameters to verify the air conditioner control performance of the Q-table and obtain the corresponding verification results; Optimizing the Q-table based on the verification results to obtain an optimized Q-table; Among them, deploying the Q-table with the optimal air-conditioning control strategy to the air-conditioning system installed on the target building, and controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy includes: Deploy the optimized Q-table with the optimal air-conditioning control strategy to the air-conditioning system installed on the target building, and control the operation of the air-conditioning system by the reinforcement learning agent based on the optimized Q-table with the optimal air-conditioning control strategy.

[0010] In one implementation, the controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy includes: Obtain the current real-time environmental data, and parse the current real-time environmental data by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy to determine the optimal air-conditioning control action, and generate an air-conditioning control instruction corresponding to the optimal air-conditioning control action; Send the air-conditioning control instruction to the infrared controller based on the Modbus protocol to control the operation of the air-conditioning system through the infrared controller.

[0011] In one implementation, during the process of controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy, it further includes: Record in real time the full-dimensional operation log including the current time, the current real-time environmental data, the air-conditioning control action, and the energy consumption of the air-conditioning system, so as to dynamically adjust the Q-table based on the full-dimensional operation log.

[0012] The present invention also discloses an air-conditioning control device based on model reinforcement learning. Among them, the device includes: An RC model building module for building an RC model reflecting the thermal behavior of the target building as a training environment model; An agent building module for building a reinforcement learning agent; An agent training module for training the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with the optimal air-conditioning control strategy; A deployment module for deploying the Q-table with the optimal air-conditioning control strategy to the air-conditioning system installed on the target building; A control module for controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy.

[0013] The present invention also discloses a terminal, which includes: a memory, a processor, and an air conditioner control program based on model reinforcement learning stored on the memory and operable on the processor. When the air conditioner control program based on model reinforcement learning is executed by the processor, the steps of the air conditioner control method based on model reinforcement learning as described above are implemented.

[0014] The present invention also discloses a computer-readable storage medium, where the computer-readable storage medium stores a computer program that can be executed to implement the steps of the air conditioner control method based on model reinforcement learning as described above.

[0015] The air conditioner control method, device, terminal, and medium based on model reinforcement learning provided by the present invention. The air conditioner control method based on model reinforcement learning includes: building an RC model reflecting the thermal behavior of a target building as a training environment model; constructing a reinforcement learning agent, and training the reinforcement learning agent in a reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioner control strategy; deploying the Q-table with the optimal air conditioner control strategy to an air conditioner system installed on the target building, and controlling the operation of the air conditioner system through the reinforcement learning agent based on the Q-table with the optimal air conditioner control strategy. It can be seen that by combining the advantages of the RC model and reinforcement learning, the present invention makes full use of the information of the interaction between reinforcement learning and the environment, reduces the demand for historical data, can train an efficient and accurate control strategy in a short time, that is, can reduce the dependence on a high-precision building physics model or a large amount of historical data, thereby realizing lightweight model construction and convenient deployment, and dynamically adjusting the air conditioner operation state through the Q-table with the optimal air conditioner control strategy deployed in the air conditioner system, taking into account user comfort and demand response objectives to the greatest extent, and realizing efficient, stable, and highly robust air conditioner demand response management. Description of the Drawings

[0016] Figure 1 is a flowchart of a preferred embodiment of the air conditioner control method based on model reinforcement learning in the present invention; Figure 2 is a schematic diagram showing the change of a specific reward function with the number of iterations disclosed in the present invention; Figure 3 is a schematic diagram showing the change of indoor temperature and power consumption in a deployment application disclosed in the present invention; Figure 4 is a schematic diagram of a reinforcement learning application disclosed in the present invention; Figure 5 is a logical block diagram of a preferred embodiment of the air conditioner control method based on model reinforcement learning in the present invention; Figure 6It is a schematic diagram showing the verification results of indoor temperature and air conditioner start-stop strategy in virtual verification disclosed by the present invention; Figure 7 It is a schematic diagram of the principle framework of an air conditioner control method based on model reinforcement learning disclosed by the present invention; Figure 8 It is a schematic diagram for comparing the results of indoor temperature changes in virtual verification and deployment applications disclosed by the present invention; Figure 9 It is a functional principle block diagram of a preferred embodiment of an air conditioner control device based on model reinforcement learning in the present invention; Figure 10 It is a functional principle block diagram of a preferred embodiment of a terminal in the present invention. Detailed implementation manners

[0017] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] Currently, several control strategies proposed for the control of air conditioner systems during demand response periods include rule-based control methods, model-based control methods, and reinforcement learning-based control methods, which are specifically as follows: (1) Rule-based control method The rule-based control method relies on pre-set rules or logics to adjust the air conditioner load. Usually, it does not require complex modeling and calculations and is applicable to scenarios with high real-time response requirements and simple control systems. When adopting the rule-based control, due to its lack of flexibility, the applicable scope and optimization degree are limited, it is difficult to cope with complex environments and cannot ensure that the system operates in an efficient state.

[0019] (2) Model-based control method The model-based control method refers to establishing a mathematical model of the air conditioner system, such as a thermodynamic model and an energy consumption model, and combining optimization algorithms to achieve refined energy-saving control. Usually, it includes model predictive control (MPC, Model Predictive Control), optimal control, and multi-objective optimal control, etc. However, the model-based method requires a large amount of historical data and sensor information to establish an accurate air conditioner model, usually lacks good robustness, is not applicable to old building complexes lacking historical data and sensors, and has problems of high development costs and complex deployment.

[0020] (3) Reinforcement learning-based control method The reinforcement learning control method automatically adjusts the control strategy through learning historical data and real-time feedback to adapt to the dynamic changes of the environment. As a feedback-based learning method, it has been widely applied to the demand response control of air conditioning systems. As a data-driven control method, reinforcement learning continuously optimizes the control strategy through interaction and learning with the environment, and can show excellent adaptability in highly complex and dynamically changing systems.

[0021] However, traditional reinforcement learning control methods have problems such as high data requirements, long training time, insufficient adaptability and long-term optimization ability, and complex model construction and deployment, making it difficult to quickly and efficiently meet the actual application requirements of demand response control. Therefore, this application provides an air conditioning control solution based on model-based reinforcement learning, which can reduce the dependence on high-precision building physical models or massive historical data, achieve lightweight model construction and convenient deployment, and dynamically adjust the operating state of the air conditioner, taking into account user comfort and demand response goals to the greatest extent, and realizing efficient, stable and highly robust air conditioning demand response management.

[0022] Please refer to Figure 1 , Figure 1 which is the flowchart of the air conditioning control method based on model-based reinforcement learning in the present invention. As shown in Figure 1 , the air conditioning control method based on model-based reinforcement learning described in the embodiments of the present invention includes: Step S11: Build an RC model reflecting the thermal behavior of the target building as the training environment model.

[0023] In this embodiment, for the target building, an RC (Resistance-Capacitance) model reflecting the indoor thermal behavior is established for the target building, that is, an RC model reflecting the thermal behavior of the target building is constructed as the training environment model.

[0024] Among them, the RC model is composed of thermal resistance and heat capacitance. Among them, 3 thermal resistance values and 3 heat capacitance values are determined by RC model training. The expression of the RC model is: ; ; ; ; ; ; ; Among them, represents the heat transfer between the indoor air and the air conditioning system, Represents the heat transfer between the building envelope and the indoor air, Represents the heat transfer between the outdoor air and the building envelope, Represents the indoor temperature, Represents the air-conditioning system temperature, Represents the building envelope temperature, Represents the outdoor temperature, Represents the air-conditioning system temperature at the Represents the air-conditioning system temperature at the Represents the indoor temperature at the Represents the indoor temperature at the Represents the building envelope temperature at the Represents the building envelope temperature at the Represents the air-conditioning system thermal resistance, Represents the indoor air thermal resistance, Represents the building envelope thermal resistance, Represents the indoor heat capacity, Represents the building heat capacity, Represents the cooling power of the air-conditioning system, Represents the indoor heat, ONOFF represents the operating state of the compressor, t represents time in seconds, h represents hours, Represents the mapping relationship between indoor heat and hours h, Represents the unique indoor heat corresponding to each hour.

[0025] Moreover, the operating state of the compressor is represented by 1 and 0, where 1 means on and 0 means off.

[0026] It should be noted that the thermal behavior of the target building (indoor) refers to the dynamic performance of a specific building in a thermal environment, including the processes of heat absorption, storage, transfer, and release, as well as the impacts of these processes on indoor temperature, energy consumption, and comfort. Its core goal is to optimize the thermal performance of the building to achieve energy conservation, environmental protection, and residential comfort. Using the RC model to replace the high-precision physical model, only a small number of parameters such as thermal resistance and heat capacity are required to construct the building thermal characteristics. Utilizing the physical characteristics of the RC model as the interaction environment for the reinforcement learning agent can reduce the dependence on high-precision building physical models or massive historical data.

[0027] Step S12: Construct a reinforcement learning agent and train the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with the optimal air-conditioning control strategy.

[0028] In this embodiment, an RC model reflecting the thermal behavior of the target building is established as the training environment model. At the same time, a reinforcement learning agent is constructed, and relevant parameters are set, including environmental parameters and basic agent parameters (i.e., Q-learning related parameters). After the construction of the RC model and the reinforcement learning agent is completed, the reinforcement learning agent is trained in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with the optimal air-conditioning control strategy. Among them, the Q-table records the values of different actions in each state for the agent to query and update. The agent represents a decision-making entity responsible for interacting with the environment to learn strategies, and finally updates the Q-table. It can be understood that in the model training stage, after the construction of the RC model and the reinforcement learning agent is completed, historical data is used to drive the agent to interact with the RC environment, and the agent is trained to learn the optimal air-conditioning control strategy and output the corresponding Q-table. This air-conditioning control strategy can improve the energy efficiency of the air-conditioning system in demand response events while ensuring indoor thermal comfort, providing an efficient and adaptive control solution for building air-conditioning demand response.

[0029] In this embodiment, constructing the reinforcement learning agent may specifically include: determining the basic agent parameters and environmental parameters, and constructing the corresponding reinforcement learning agent according to the basic agent parameters and environmental parameters; among them, the basic agent parameters include the state space, action space, reward function, and ε-greedy policy setting value. The ε-greedy policy setting value is a constant between 0 and 1, and the environmental parameters include outdoor temperature, indoor temperature, air-conditioning operating status, electricity price, and time information; And the Q-value update formula of the reinforcement learning agent is: ; where s represents the state space, and the state space includes indoor temperature, outdoor temperature, and time information, represents the next state space after performing action a in state s, a represents the action space, and the action space is a binary discrete action set {0, 1}, where the element 0 represents the air-conditioning is turned off, and the element 1 represents the air-conditioning is turned on, represents the action space corresponding to the next state space r represents the reward function, and the reward function consists of temperature comfort penalty, electricity cost penalty, low-temperature reward, high-temperature reward, and pre-cooling reward, represents the learning rate, represents the discount factor.

[0030] It should be noted that the ε-greedy strategy selects actions based on Q-values, that is, a constant ε between 0 and 1 is set, and the action with the largest Q-value is selected with a probability of 1 - ε as the action at the current moment, and an action is randomly selected with a probability of ε. Moreover, the above reward function r is set with the optimization goal of reducing energy costs while ensuring indoor thermal comfort. The reward function r consists of temperature comfort penalty (A), electricity cost penalty (B), low temperature reward (C), high temperature reward (D), and pre-cooling reward (E), that is: ; Among them, the temperature comfort penalty (A): includes the penalty generated due to the indoor temperature exceeding the set comfort range, such as being lower than the comfort lower limit or higher than the comfort upper limit. In the non-pre-cooling stage, if the temperature crosses the boundary, a larger penalty is imposed; while in the transition period, such as 30 minutes before 12:00 and 30 minutes before 17:00, it is appropriately relaxed to avoid excessive penalty due to too rapid temperature fluctuations. The electricity cost penalty (B): when the air conditioner is in the on state (action = 1), the corresponding penalty is calculated according to the current time-of-use electricity price and the preset cost multiplier, aiming to suppress frequent start-stop and reduce electricity expenses. The low temperature reward (C): when the air conditioner is off (action = 0), if the indoor temperature does not exceed the comfort upper limit of the current period, a positive incentive is given to encourage reducing energy consumption while meeting comfort requirements. The high temperature reward (D): within the peak period, if the indoor temperature is maintained between 26.5°C and 27°C, the system gives an additional reward. This helps to make full use of the regulation ability of the equipment near the thermal comfort boundary, thereby reducing the energy consumption waste caused by frequent start-stop. The pre-cooling reward (E): for load management before the power peak, a pre-cooling stage is set up, such as after 10:45 or after 14:45. In the pre-cooling stage, if the indoor temperature is maintained between 23.5°C and 24.5°C, a positive incentive is given; otherwise, if the temperature does not reach the expected target, a penalty is imposed to prompt the air conditioning system to achieve the expected temperature as soon as possible, creating favorable conditions for load regulation during the peak period.

[0031] In this embodiment, the reinforcement learning agent is trained in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with the optimal air conditioning control strategy, which specifically includes: collecting the historical outdoor temperature data and historical electricity price data of the target building; using the historical outdoor temperature data and historical electricity price data to drive the reinforcement learning agent to interact with the reinforcement learning simulation environment generated based on the RC model, so as to train the reinforcement learning agent to select actions based on the states obtained from the RC model, receive the rewards feedback by the environment, and update the Q-table until the training converges to obtain a Q-table with the optimal air conditioning control strategy.

[0032] For example, during the training phase of the reinforcement learning agent, the agent is trained based on the RC model, using data such as outdoor temperature, indoor temperature, air conditioner operating status (ON, OFF), electricity price, and time as inputs. First, the agent's observations and environmental parameters are determined, the RC model parameters are loaded, and the multi-day historical outdoor temperature data and historical electricity price data collected are obtained to simulate real environmental conditions (i.e., the reinforcement learning simulation environment). Second, the hyperparameter values related to Q-learning are set. Finally, the change in the indoor temperature of the building is transmitted through the RC model, the agent obtains the current state quantity, and updates the value function according to the current state quantity, and finally obtains the optimal action value function (the Q-table with the optimal air conditioner control strategy).

[0033] For another example, the training period can be set from 8:00 to 18:00 every day. The agent makes an action decision every 1 minute, that is, it executes 600 actions every day, and accumulatively generates 600 state transitions (a four-tuple process including state, action, reward, and new state). The complete interaction process of the agent from 8:00 to 18:00 every day is called an Episode. After each Episode ends, the system records the actual return of this Episode and calculates the average return of the most recent n Episodes to intuitively reflect the training process, where n can be set to 600. During the training process, the reward function changes with the number of iterations, as shown in Figure 2 shown. Among them, the operating period and temperature range are clearly divided: the peak period is defined as 11:00–12:00 and 15:00–17:00, totaling 3 hours. Since the electricity cost is high during this period, the system allows a wider temperature range, such as 23°C to 27°C, in order to perform flexible regulation on the premise of meeting the basic comfort requirements. For the non-peak period, to ensure more strict control of the indoor temperature fluctuation, the comfortable temperature range is set to 24°C to 26°C.

[0034] Step S13: Deploy the Q-table with the optimal air conditioner control strategy to the air conditioner system installed on the target building, and control the operation of the air conditioner system through the reinforcement learning agent based on the Q-table with the optimal air conditioner control strategy.

[0035] In this embodiment, after the agent completes training and outputs the Q-table with the optimal air conditioner control strategy, it can be deployed and applied, that is, the Q-table is deployed to the air conditioner system installed on the target building, and the operation of the air conditioner system is controlled through the reinforcement learning agent based on the Q-table with the optimal air conditioner control strategy. It can be understood that when applying the control strategy to the actual operation of the building air conditioner system and performing real-time indoor and outdoor temperature monitoring, the agent analyzes the indoor and outdoor temperatures based on the control strategy and outputs control actions. Among them, the changes in indoor temperature and power consumption are shown in Figure 3 shown.

[0036] Specifically, obtain the current real-time environmental data, and use the reinforcement learning agent to analyze the current real-time environmental data based on the Q-table with the optimal air-conditioning control strategy to determine the optimal air-conditioning control action, and generate an air-conditioning control instruction corresponding to the optimal air-conditioning control action; send the air-conditioning control instruction to the infrared controller based on the Modbus protocol to control the operation of the air-conditioning system through the infrared controller. It should be noted that at a high time granularity, using the Modbus protocol to ensure the real-time data transmission, in response to sudden weather changes and equipment performance changes, it can quickly respond and dynamically adjust to ensure that it can adapt to the actual building environment and achieve minute-level control.

[0037] In this embodiment, during the process of controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy, it further includes: real-time recording of the full-dimensional operation log including the current time, the current real-time environmental data, the air-conditioning control action, and the energy consumption of the air-conditioning system, so as to dynamically adjust the Q-table based on the full-dimensional operation log.

[0038] For example, as shown in Figure 4 First, at the hardware level, the start and stop of the air-conditioning system are controlled by the infrared controller. The infrared controller is connected to the Python console through the Modbus protocol, which can achieve fast and reliable data interaction, including sending start and stop instructions to the air-conditioning and obtaining the operation status feedback from the device side. At the same time, the currently obtained real-time environmental data, such as the indoor and outdoor temperature data of the temperature and humidity sensor and the outdoor weather station, as well as the external time-varying electricity price information, are used as the input of the agent to form the main source of the environmental state. Specifically, through the wireless network (Wi-Fi or 4G), the Python console can obtain the above environmental data in real time and refresh the agent's perception of the environment at short time intervals; secondly, at the software level, the agent and the Python console together constitute the core of decision-making and execution. The agent infers the environmental state collected each time based on the trained Q-table, so as to select the optimal action, that is, to decide whether to turn on or off the air-conditioning; finally, the Python console sends the action instruction to the infrared controller in the form of the Modbus protocol, which triggers the air-conditioning operation to achieve direct interaction with the physical device. In order to realize the online or offline update of reinforcement learning, the corresponding immediate reward can be calculated after each action is executed, and the Q-table can be adjusted online when necessary, so as to maintain the adaptability between the policy and the dynamic changes of the environment.

[0039] That is to say, first, it is deployed at the hardware interaction layer. The two-way communication between the infrared controller and the Python console is realized through the Modbus protocol to issue air conditioner control instructions in real time and obtain device status feedback. Secondly, it is implemented at the software decision-making layer. The intelligent agent infers based on the Q-table for real-time environmental data (temperature and electricity price) to generate air conditioner control action instructions. Finally, a data closed-loop and policy adjustment are formed. In the actual air conditioner system control process, a full-dimensional operation log including the current time, current real-time environmental data, air conditioner control actions, and the energy consumption of the air conditioner system can be established. By analyzing the environmental changes offline and automatically correcting the policy parameters, the comfort and energy efficiency are continuously balanced. That is to say, the reward value is calculated after each action is executed, and the Q-table is adjusted in real time when necessary, taking into account the user's comfort and demand response goals to the greatest extent, and realizing efficient, stable, and highly robust air conditioner demand response management.

[0040] It can be seen that in the embodiment of the present invention, by combining the advantages of the RC model and reinforcement learning, fully utilizing the information of the interaction between reinforcement learning and the environment, the requirement for historical data is reduced, and an efficient and accurate control strategy can be trained in a short time. That is to say, the dependence on a high-precision building physics model or a large amount of historical data can be reduced, so as to realize lightweight model construction and convenient deployment. And by deploying the Q-table with the optimal air conditioner control strategy in the air conditioner system, the air conditioner operation state is dynamically adjusted, taking into account the user's comfort and demand response goals to the greatest extent, and realizing efficient, stable, and highly robust air conditioner demand response management.

[0041] That is to say, the air conditioner control method based on model reinforcement learning in this application combines the advantages of the RC model and reinforcement learning, uses the RC model to replace the high-precision physical model, and only needs a small number of parameters such as thermal resistance and heat capacity to construct the building thermal characteristics. Using the physical characteristics of the RC model as the interaction environment of the reinforcement learning intelligent agent reduces the dependence on a high-precision building physical model or a large amount of historical data. Through the real-time interaction between the Q-learning algorithm and the RC model, the construction of the RC model can be completed in only one day, and the training of the reinforcement learning Q-table can be completed within 10 minutes. That is to say, the air conditioner control method based on model reinforcement learning in this application is a lightweight and efficient air conditioner control scheme with less data requirements, fast training speed, and strong adaptability. It has the ability of long-term optimization and hourly deployment, can overcome the problems of high data requirements, long training time, insufficient adaptability and long-term optimization ability, and complex model construction and deployment existing in the traditional air conditioner system control method, and improves the robustness and optimization ability of the system. Compared with the problems existing in the traditional air conditioner control method, the air conditioner control method based on model reinforcement learning in this application has strong scalability and strong control strategy adaptability, effectively solves the actual application requirements that the traditional method is difficult to meet the demand response control quickly and efficiently, and provides an efficient and stable solution for the demand response control.

[0042] See Figure 5 As shown, an embodiment of the present invention discloses a specific air conditioner control method based on model reinforcement learning. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution.

[0043] Step S21: Build an RC model reflecting the thermal behavior of the target building as the training environment model.

[0044] Step S22: Construct a reinforcement learning agent and train the reinforcement learning agent in a reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioner control strategy.

[0045] Step S23: Deploy the Q-table with the optimal air conditioner control strategy in the reinforcement learning simulation environment generated based on the RC model with the same training environment parameters to verify the air conditioner control performance of the Q-table and obtain the corresponding verification result.

[0046] In this embodiment, before deploying and applying the air conditioner control strategy, it can be verified. The Q-table with the optimal air conditioner control strategy output by the agent is deployed in a reinforcement learning simulation environment generated based on the RC model with the same training environment parameters to verify the actual control performance through a virtual interaction environment constructed with the target building, and the corresponding verification result is obtained. It can be understood that the effectiveness and stability of the model strategy are verified on the RC model, and the Q-table output by the agent is deployed on the RC model for verification to evaluate the control strategy performance.

[0047] For example, during the verification process, first, load the Q-table and RC model parameters with the optimal air-conditioning control strategy generated in the training stage to ensure that the same environmental settings as those during training are used in the verification process. It is also necessary to obtain new outdoor temperature data to simulate the actual temperature changes. Secondly, the agent selects the optimal action based on the current state (indoor and outdoor temperatures, time) and the Q-table. The environmental changes per minute are used to calculate the changes in indoor temperature, system temperature, and wall temperature through the state transition function. Record the action selection and temperature changes in real time to simulate the performance of the control strategy in the actual environment. That is, input the newly collected outdoor temperature data into the agent, and the agent selects the optimal action based on the current state and the optimal action value function. Among them, the changes in indoor temperature, system temperature, and wall temperature can be calculated through the RC model, and the action selection and temperature changes are recorded in real time, and the immediate reward is calculated and the state changes, control actions, and reward values at each time are recorded. Based on the logged data recorded in real time, draw detailed temperature change curves, air-conditioning start / stop states, and reward curves to comprehensively evaluate the overall effect and performance stability of the strategy. Finally, visualize the verification results through a time series graph, including indoor temperature, outdoor temperature, air-conditioning actions, and immediate reward curves. By comparing the actual temperature changes and the performance of the control strategy, the performance of the Q-learning model in different scenarios can be analyzed in depth, and the effectiveness and robustness of the strategy can be further improved.

[0048] Among them, during the verification stage, three groups of typical weather data with significant temperature differences can be selected: the average temperature on the first day is 28.8 °C (mild working condition), the average temperature on the second day is 33.4 °C (high temperature working condition), and the average temperature on the third day is 32.4 °C (fluctuating working condition). The environmental adaptability of the Q-table is investigated through a multi-condition comparison test system. The verification results of one day are shown in Figure 6 as follows. The analysis of the verification results is as follows: First, in terms of temperature control, the system successfully maintains the room temperature within the 26 °C threshold during non-demand response periods and effectively controls it below the 27 °C limit during demand response periods. Before the DR events (11-12 o'clock, 15-17 o'clock) on the three verification days, the agent always activates the precooling strategy in advance, that is, the temperature is reduced in advance by 1.2 - 1.5 °C; Second, in terms of air-conditioning operation regulation, the air-conditioning operation duration during the first round of DR periods is reduced by 92.3% compared with the baseline condition, that is, it can be reduced from the conventional 54.5 minutes to 4.2 minutes; during the second round of DR periods, the reduction rates can reach 79.7%, 75.5%, and 75.7% respectively. The electricity consumption reduction rates of the three groups all exceed 75%.

[0049] Step S24: Optimize the Q-table based on the verification results to obtain an optimized Q-table.

[0050] In this embodiment, after verifying the effectiveness and stability of the model strategy on the RC model, the Q-table can be optimized based on the verification results to obtain an optimized Q-table. It can be understood that when it is found during the verification stage that the effectiveness and stability of the strategy do not meet the requirements, the Q-table can be further optimized so that efficient and accurate control can be achieved in subsequent deployment applications.

[0051] Step S25: Deploy the optimized Q-table with the optimal air-conditioning control strategy to the air-conditioning system installed on the target building, and control the operation of the air-conditioning system by the reinforcement learning agent based on the optimized Q-table with the optimal air-conditioning control strategy.

[0052] For the specific content of the above steps S21 to S22 and step S25, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0053] It can be seen that in the embodiment of the present invention, by combining the advantages of the RC model and reinforcement learning, making full use of the information of the interaction between reinforcement learning and the environment, the demand for historical data is reduced, and an efficient and accurate control strategy can be trained in a short time. That is, the dependence on a high-precision building physics model or a large amount of historical data can be reduced, so as to realize lightweight model construction and convenient deployment. And by deploying the Q-table with the optimal air-conditioning control strategy in the air-conditioning system, the operation state of the air-conditioning is dynamically adjusted, taking into account the user comfort and the demand response target to the greatest extent, and realizing efficient, stable and highly robust air-conditioning demand response management. The technical solution of this application forms an adaptive closed-loop control system from the early-stage simulation and verification to the later-stage field deployment, which can dynamically adjust the operation state of the air-conditioning, taking into account the user comfort and the demand response target to the greatest extent, and realizing efficient, stable and highly robust air-conditioning demand response management.

[0054] For example, see Figure 7As shown in the figure, the present application combines the advantages of the RC model and reinforcement learning to propose a three-step strategy of "model training + virtual verification + deployment and application" to achieve a lightweight and efficient reinforcement learning control method. That is, in model training, an RC model reflecting the indoor thermal behavior of the target building is established, and a reinforcement learning agent is constructed, and relevant parameters are set, such as the basic parameters of the agent including the state space, action space, reward function, and ε-greedy policy setting value, and the environmental parameters including outdoor temperature, indoor temperature, air-conditioning operation status, electricity price, and time information. Then, historical data is used to drive the interaction between the agent and the RC environment to output the Q-table after training is completed. In virtual verification, the effectiveness and stability of the model strategy are verified on the RC model. The Q-table output by the agent is deployed on the RC model for verification to evaluate the performance of the control strategy. In deployment and application, the control strategy is applied to the operation of the actual building air-conditioning system, and real-time indoor and outdoor temperature monitoring is carried out. The agent analyzes the indoor and outdoor temperatures based on the control strategy and outputs control actions. This solution of the present application can solve the problems of complex modeling, poor policy flexibility, high deployment cost, and complex deployment in traditional control methods, and achieve double optimization of building thermal comfort and energy efficiency by dynamically adjusting the air-conditioning control strategy. And, for the comparison of the results of the indoor temperature change in virtual verification and deployment and application, please refer to Figure 8 As shown in the figure, the outdoor temperature data of the experimental day is imported into the RC model to obtain a "virtual verification - indoor temperature" curve obtained by the virtual simulation method. There are certain differences between the two curves in local time periods, mainly due to the different thermodynamic characteristics and external interference factors between the experimental environment and the virtual simulation environment. However, from the overall trend, the two show a highly consistent fluctuation law, and the correlation coefficient for the whole day can reach 0.87, indicating that the virtual simulation can well simulate the key features in the experimental process at the macroscopic level.

[0055] In one embodiment, as Figure 9 As shown in the figure, based on the above air-conditioning control method based on model reinforcement learning, the present invention also correspondingly provides an air-conditioning control device based on model reinforcement learning, including: An RC model building module 11, configured to build an RC model reflecting the thermal behavior of the target building as a training environment model; An agent construction module 12, configured to construct a reinforcement learning agent; An agent training module 13, configured to train the reinforcement learning agent in a reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air-conditioning control strategy; A deployment module 14, configured to deploy the Q-table with the optimal air-conditioning control strategy to the air-conditioning system carried on the target building; A control module 15, configured to control the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy.

[0056] Figure 10 The following is a schematic structural diagram of the terminal provided by the embodiment of the present application. The terminal may include: A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0057] When the processor 502 executes the program, it implements the air-conditioning control method based on model-based reinforcement learning provided in the above embodiment.

[0058] Further, the terminal further includes: A communication interface 503, configured for communication between the memory 501 and the processor 502.

[0059] The memory 501 is used to store a computer program executable on the processor 502.

[0060] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0061] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be interconnected through a bus and complete communication with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0062] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a chip, the memory 501, the processor 502, and the communication interface 503 may complete communication with each other through an internal interface.

[0063] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0064] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-described air conditioner control method based on model reinforcement learning is implemented.

[0065] Those skilled in the art will readily conceive of other implementations of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0066] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0067] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can read instructions from the instruction execution system, apparatus, or device and execute the instructions.

[0068] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0069] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or changes can be made according to the above description, and all such improvements and changes should fall within the protection scope of the appended claims of the present invention.

Claims

1. An air conditioner control method based on model reinforcement learning, characterized in that The method includes: Construct an RC model reflecting the thermal behavior of the target building as the training environment model; Construct a reinforcement learning agent and train the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioning control strategy; Deploy the Q-table with the optimal air conditioning control strategy to the air conditioning system installed on the target building, and control the operation of the air conditioning system through the reinforcement learning agent based on the Q-table with the optimal air conditioning control strategy.

2. The air conditioner control method based on model reinforcement learning according to claim 1, characterized in that, The RC model is composed of thermal resistance and heat capacity, and the expression of the RC model is: ; ; ; ; ; ; ; Among them, represents the heat transfer between the indoor air and the air conditioning system, represents the heat transfer between the building envelope and the indoor air, represents the heat transfer between the outdoor air and the building envelope, represents the indoor temperature, represents the air conditioning system temperature, represents the building envelope temperature, represents the outdoor temperature, represents the air conditioning system temperature at the represents the air conditioning system temperature at the represents the indoor temperature at the represents the indoor temperature at the represents the building envelope temperature at the represents the building envelope temperature at the represents the air conditioning system thermal resistance, represents the indoor air thermal resistance, represents the building envelope thermal resistance, represents the indoor heat capacity, represents the building heat capacity, represents the cooling power of the air conditioning system, represents the indoor heat, ONOFF represents the operating state of the compressor, t represents time in seconds, h represents hours, represents the mapping relationship between the indoor heat and hours h, represents the indoor heat corresponding to each hour.

3. The air conditioner control method based on model reinforcement learning according to claim 1, wherein, The construction of the reinforcement learning agent includes: Determine the basic parameters of the agent and the environmental parameters, and construct the corresponding reinforcement learning agent according to the basic parameters of the agent and the environmental parameters; Among them, the basic parameters of the agent include the state space, action space, reward function, and ε-greedy policy setting value. The ε-greedy policy setting value is a constant between 0 and 1. The environmental parameters include outdoor temperature, indoor temperature, air conditioning operation status, electricity price, and time information; And the Q-value update formula of the reinforcement learning agent is: ; Among them, s represents the state space, and the state space includes indoor temperature, outdoor temperature, and time information. represents the next state space after performing action a in state s, a represents the action space, and the action space is a binary discrete action set {0, 1}, where element 0 represents the air conditioner is turned off, and element 1 represents the air conditioner is turned on. represents in the next state space the corresponding action space, r represents the reward function, and the reward function consists of temperature comfort penalty, electricity cost penalty, low temperature reward, high temperature reward, and pre-cooling reward. represents the learning rate. represents the discount factor.

4. The air conditioner control method based on model reinforcement learning according to claim 1, wherein The training of the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with an optimal air conditioning control strategy includes: Collect the historical outdoor temperature data and historical electricity price data of the target building; Use the historical outdoor temperature data and the historical electricity price data to drive the interaction between the reinforcement learning agent and the reinforcement learning simulation environment generated based on the RC model, so as to train the reinforcement learning agent to select actions based on the states obtained from the RC model, receive the rewards feedback by the environment, and update the Q-table until the training converges to obtain a Q-table with an optimal air conditioning control strategy.

5. The air conditioner control method based on model reinforcement learning according to claim 1, wherein Before deploying the Q-table with the optimal air conditioning control strategy to the air conditioning system installed on the target building and controlling the operation of the air conditioning system through the reinforcement learning agent based on the Q-table with the optimal air conditioning control strategy, it further includes: Deploy the Q-table with the optimal air conditioning control strategy to the reinforcement learning simulation environment generated based on the RC model with the same training environment parameters to verify the air conditioning control performance of the Q-table and obtain the corresponding verification results; Optimize the Q-table based on the verification results to obtain an optimized Q-table; Among them, the deployment of the Q-table with the optimal air conditioning control strategy to the air conditioning system installed on the target building and the control of the operation of the air conditioning system through the reinforcement learning agent based on the Q-table with the optimal air conditioning control strategy include: Deploy the optimized Q-table with the optimal air conditioning control strategy to the air conditioning system installed on the target building, and control the operation of the air conditioning system through the reinforcement learning agent based on the optimized Q-table with the optimal air conditioning control strategy.

6. The air conditioner control method based on model reinforcement learning according to any one of claims 1 to 5, characterized in that The control of the operation of the air conditioning system through the reinforcement learning agent based on the Q-table with the optimal air conditioning control strategy includes: Obtain the current real-time environmental data, and parse the current real-time environmental data by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy to determine the optimal air-conditioning control action, and generate an air-conditioning control instruction corresponding to the optimal air-conditioning control action; Send the air-conditioning control instruction to the infrared controller based on the Modbus protocol to control the operation of the air-conditioning system through the infrared controller.

7. The air conditioner control method based on model reinforcement learning according to claim 6, characterized in that During the process of controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy, it further includes: Record in real time the full-dimensional operation log including the current time, the current real-time environmental data, the air-conditioning control action and the energy consumption of the air-conditioning system, so as to dynamically adjust the Q-table based on the full-dimensional operation log.

8. An air conditioner control device based on model reinforcement learning, characterized in that, The device includes: An RC model building module for building an RC model reflecting the thermal behavior of the target building as a training environment model; An agent building module for building a reinforcement learning agent; An agent training module for training the reinforcement learning agent in the reinforcement learning simulation environment generated based on the RC model to obtain a Q-table with the optimal air-conditioning control strategy; A deployment module for deploying the Q-table with the optimal air-conditioning control strategy to the air-conditioning system carried on the target building; A control module for controlling the operation of the air-conditioning system by the reinforcement learning agent based on the Q-table with the optimal air-conditioning control strategy.

9. A terminal, characterized in that, It includes: A memory, a processor, and an air-conditioning control program based on model reinforcement learning stored on the memory and executable on the processor. When the air-conditioning control program based on model reinforcement learning is executed by the processor, it implements the steps of the air-conditioning control method based on model reinforcement learning according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the air-conditioning control method based on model reinforcement learning according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Reinforcement learning modeling method for demand response of building air conditioning system

    CN113435042A

  • Indoor thermal environment learning efficiency improvement optimization control method based on reinforcement learning

    CN114370698A

  • Heating ventilation air conditioner regulation and control method and device based on reinforcement learning

    CN115950080A

  • Indoor thermal environment control method based on RC model and deep reinforcement learning

    CN116734424A