Air conditioner control method, machine readable storage medium and air conditioner
By optimizing the control decision model of the central air-conditioning system through reinforcement learning algorithms, the energy consumption problem of the central air-conditioning system while meeting the comfort and stability of the user end is solved, achieving energy conservation and emission reduction and improving system stability.
Patent Information
- Application Number
- CN202410343919.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing central air-conditioning systems are difficult to effectively save energy while meeting user comfort and system stability, and the control system is prone to violent fluctuations.
A control decision model based on reinforcement learning algorithm is adopted. The decision parameters are obtained through Actor-Critic network training. The rated power utilization rate of the outdoor unit and the indoor temperature are taken as the adjustment targets to optimize the working state of the compressor.
Under the premise of meeting user-side comfort and system stability, it effectively saves outdoor unit energy consumption, improves system stability and reduces energy waste.
Smart Images

Figure CN120702074A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air treatment equipment, and in particular to a control method for an air conditioner, a machine-readable storage medium, and an air conditioner. Background Art
[0002] In the field of air handling equipment, for example, central air conditioning systems typically utilize a single outdoor unit to provide heating and cooling for multiple user-end devices. Because user demand is volatile and unpredictable, minimizing energy consumption in the outdoor unit while ensuring the comfort of each user and system operational stability has become a pressing challenge.
[0003] To address this issue, existing technologies have proposed using actual system status parameters to characterize changes in user demand, thereby adjusting the operating conditions of each outdoor unit's compressors, such as frequency and number of units powered on, to meet user comfort requirements. However, this technical solution can only respond passively to user demand. On the one hand, it fails to fully tap the energy-saving potential of the central air conditioning system, and energy consumption remains high. On the other hand, when users make short-term adjustments, the control system will passively adjust for a short period of time, causing significant fluctuations in the entire system and affecting system stability. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed to provide a control method, machine-readable storage medium and air conditioner for an air conditioner that overcome the above problems or at least partially solve the above problems, aiming to save energy consumption of the outdoor unit and achieve the purpose of energy conservation and emission reduction while satisfying the comfort of each user end and the stability of the system.
[0005] Specifically, the present invention provides the following technical solutions:
[0006] A method for controlling an air conditioner, comprising:
[0007] Obtaining the operating mode and system status parameters of the air conditioner;
[0008] Obtain a control decision model based on reinforcement learning algorithm;
[0009] Obtaining the decision parameter according to the input parameter and the control decision model;
[0010] A control action is determined according to the decision parameter to control the air conditioner.
[0011] Wherein, the control decision model is configured to: take the working mode and the system state parameters as input parameters, at least take the rated power utilization rate and indoor temperature of the outdoor unit of the air conditioner as adjustment targets, and output decision parameters for adjusting the outdoor unit.
[0012] Optionally, the outdoor unit comprises at least one compressor, and each compressor is configured to adjust its working state according to at least a high pressure preset value of the outdoor unit or a low pressure preset value of the outdoor unit; and
[0013] The working modes include cooling mode and heating mode.
[0014] Optionally, the system state parameters include state parameters and time parameters, the state parameters include outdoor temperature, indoor temperature, frequency of each compressor, power of each compressor, high pressure of each compressor, low pressure of each compressor, heating capacity of the outdoor unit, high pressure preset value, low pressure preset value, and the time parameters include time tags of each state parameter.
[0015] Optionally, the decision parameters include a change in a high-pressure preset value in the heating mode and a change in a low-pressure preset value in the cooling mode.
[0016] Optionally, obtaining a control decision model based on a reinforcement learning algorithm includes:
[0017] Build a reinforcement learning model based on the Actor-Critic network;
[0018] With a first preset time length as a first cycle, a cyclic interaction is established between the air conditioner and the reinforcement learning model to perform online iterative training on the reinforcement learning model and obtain the control decision model.
[0019] Optionally, before inputting the input parameters into the reinforcement learning model for online training, the method further includes:
[0020] acquiring the input parameters of the air conditioner in each of the first cycles;
[0021] The system state parameters of each of the first periods are integrated.
[0022] The integration process includes:
[0023] Performing interpolation processing on abnormal parameters in the system state parameters to obtain usable system state parameters; the abnormal parameters include one or more of missing parameters, parameters not within a preset threshold range, and outlier numerical parameters;
[0024] performing time alignment on the state parameters in the available system state parameters according to the time parameters;
[0025] Resampling is performed on each of the state parameters with the first preset time length as a step length.
[0026] Optionally, the construction of a reinforcement learning model based on an Actor-Critic network includes:
[0027] Constructing an interactive environment based on the air conditioner;
[0028] Constructing a return pool to store the input parameters and the decision parameters;
[0029] Build an Actor-Critic network agent.
[0030] Optionally, the construction is based on the interactive environment of the air conditioner, including:
[0031] Setting environmental state parameters, wherein the environmental state parameters include the input parameters;
[0032] Setting control action parameters, wherein the control action parameters include the decision parameters;
[0033] A reward and penalty function is constructed, where the reward and penalty function is configured to be negatively correlated with at least the rated power utilization of the outdoor unit and negatively correlated with the difference between the indoor temperature and the desired indoor temperature.
[0034] Optionally, the constructing and returning to the pool comprises:
[0035] An environmental data database is constructed. The environmental database is configured to persistently store a data set generated each time the intelligent agent interacts with the air conditioner. The data set includes at least the environmental state parameters, the control action parameters, and the result of the reward and punishment function.
[0036] Optionally, the construction of the Actor-Critic network agent includes:
[0037] Build the Actor initial network and Critic network;
[0038] Acquire pre-training sample data, wherein the pre-training sample data at least includes the operating mode of the air conditioner, the system state parameters, and historical data of the control action;
[0039] The Actor initial network is pre-trained according to the pre-training sample data to obtain an Actor network.
[0040] Optionally, the system status parameters further include outdoor humidity and indoor humidity; and
[0041] The reward and penalty function is further configured to be negatively correlated with a difference between the indoor humidity and a desired indoor humidity.
[0042] Optionally, the constructing of the reward and punishment function includes:
[0043] Obtaining a preset threshold value of the system status parameter;
[0044] The reward and punishment function is configured to be at least negatively correlated with the difference between the system state parameter and the preset threshold.
[0045] On the other hand, the present invention further provides a machine-readable storage medium having a machine-executable program stored thereon. When the machine-executable program is executed by a processor, the air conditioner control method as described in any one of the above items is implemented.
[0046] On the other hand, the present invention also provides an air conditioner, which includes a controller, the controller including a memory, a processor and a machine executable program stored in the memory and running on the processor, and when the processor executes the machine executable program, it implements the control method of the air conditioner as described in any one of the above items.
[0047] Optionally, the air conditioner is a multi-split air conditioner, wherein the outdoor unit of the multi-split air conditioner comprises a plurality of compressors, and each compressor is configured to adjust its working state according to at least a high pressure preset value of the outdoor unit or a low pressure preset value of the outdoor unit.
[0048] The air conditioner control method, machine-readable storage medium, and air conditioner of the present invention include a control decision model based on a reinforcement learning algorithm. The control decision model can output decision parameters for adjusting the outdoor unit based on the operating mode and system state parameters. The control decision model is trained using a reinforcement learning algorithm and configured to use the operating mode and system state parameters as input parameters, with at least the rated power utilization rate of the air conditioner's outdoor unit and the indoor temperature as adjustment targets. The control decision model can fully leverage the advantages of the reinforcement learning algorithm in sequential decision-making tasks, deeply exploring the inherent relationship between user-end requirements, system decision control, and decision control results. This can thereby achieve energy savings for the outdoor unit while satisfying the comfort requirements of each user, thereby achieving the goal of energy conservation and emission reduction.
[0049] On the other hand, the control method of the present application, since the reinforcement learning algorithm accumulates rewards over a longer period, will take into account the changes in user-end load and control system over a longer period of time when outputting decision parameters based on input parameters. While further saving energy consumption of the outdoor unit, it can also improve the stability of the system.
[0050] Furthermore, the control method of this application, when constructing an actor-critic network agent, first obtains pre-training sample data to pre-train the initial actor network. This allows the actor-critic network agent to output relatively reasonable decision parameters during online training of the reinforcement learning model, avoiding situations in which the actor-critic network agent randomly outputs decision parameters that fail to meet user requirements in the early stages of training, thereby impacting user comfort.
[0051] Furthermore, the control method of the present application, when training a reinforcement learning model, persistently stores the data sets generated each time the agent interacts with the air conditioner in a return pool. This allows the data in the return pool to be reused during subsequent training to adjust the reward and penalty functions, train the critic network, and pre-train the actor network. This improves the utilization of the data in the return pool and enhances the efficiency of reinforcement learning training.
[0052] Furthermore, in the control method of the present application, when training the reinforcement learning model, the reward and penalty function is configured to be at least negatively correlated with the difference between the system state parameter and the preset threshold. This ensures that the air conditioning system operates within its normal operating range when outputting the decision parameters, avoiding drastic fluctuations that could affect overall system stability. This reduces or prevents outdoor unit downtime and improves system stability.
[0053] Therefore, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Hereinafter, some specific embodiments of the present invention will be described in detail in an exemplary and non-limiting manner with reference to the accompanying drawings. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the accompanying drawings:
[0055] Figure 1 is a schematic flow chart of a method for controlling an air conditioner according to an embodiment of the present invention;
[0056] Figure 2 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0057] Figure 3 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0058] Figure 4 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0059] Figure 5 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0060] Figure 6 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0061] Figure 7 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0062] Figure 8 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0063] Figure 9 is a schematic flow chart of a control method according to one embodiment of the present invention;
[0064] Figure 10 is a schematic block diagram of a control method according to an embodiment of the present invention;
[0065] Figure 11 is a schematic block diagram of a machine-readable storage medium according to one embodiment of the present invention;
[0066] Figure 12 is a schematic block diagram of an air conditioner according to an embodiment of the present invention. DETAILED DESCRIPTION
[0067] Refer to the following Figures 1 to 12 The present invention will be described in detail with reference to an embodiment of the present invention to describe a method for controlling an air conditioner, a machine-readable storage medium, and an air conditioner. The terms "front," "rear," "upper," "lower," "top," "bottom," "inner," "outer," and "lateral" are used to indicate directions or positions based on those shown in the accompanying drawings. These terms are intended solely to facilitate and simplify the description of the present invention and do not indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific direction. Therefore, they should not be construed as limiting the present invention.
[0068] The terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the definition of "first", "second", etc. can explicitly or implicitly include at least one of the features, that is, include one or more of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. When a feature "includes or contains" one or more of the features it covers, unless otherwise specifically described, this indicates that other features are not excluded and may further include other features.
[0069] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," "fixed," "coupled," and the like should be interpreted broadly. For example, they may refer to fixed or detachable connections, or integration; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components or interaction between two components, unless otherwise specified. A person of ordinary skill in the art should be able to understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0070] Figure 1 is a schematic flow chart of a method for controlling an air conditioner according to an embodiment of the present invention, and Figure 2-12 , the present invention provides a control method for an air conditioner.
[0071] A method for controlling an air conditioner comprises the following steps:
[0072] S100, obtaining the operating mode and system status parameters of the air conditioner;
[0073] S200, obtaining a control decision model based on a reinforcement learning algorithm;
[0074] S300, obtaining decision parameters according to the input parameters and the control decision model;
[0075] S400: Determine a control action according to the decision parameters to control the air conditioner.
[0076] The control decision model is configured to take the working mode and system state parameters as input parameters, at least the rated power utilization rate and indoor temperature of the outdoor unit of the air conditioner as adjustment targets, and output decision parameters for adjusting the outdoor unit.
[0077] In this embodiment, the operating mode of the air conditioner can be heating mode, cooling mode, dehumidification mode, humidification mode, air supply mode, etc. In different operating modes, the system state parameters of the air conditioner may be different and affect the control action of the outdoor unit.
[0078] System status parameters may include outdoor temperature, indoor temperature, frequency of each compressor, power of each compressor, high pressure of each compressor, low pressure of each compressor, heating capacity of the outdoor unit, preset high pressure value, preset low pressure value, outdoor humidity, indoor humidity, etc. These parameters can be detected by multiple sensors. System status parameters may also include time stamps for the detection of these parameters. For each parameter type, its time stamp can be combined to form a corresponding time series. Indoor temperature and indoor humidity may be average indoor temperatures, such as the arithmetic mean or weighted mean of the temperatures of each room.
[0079] The rated power utilization rate of an outdoor unit can be calculated as the ratio of the current power to the rated power. The rated power is the maximum power of the outdoor unit under normal operating conditions. While meeting user comfort requirements, the lower the rated power utilization rate, the lower the outdoor unit's energy consumption, resulting in greater energy savings and emissions reductions.
[0080] In this embodiment, the control decision model uses indoor temperature as the adjustment target for decision-making to meet the user's temperature control needs. In other embodiments of the present invention, the control decision model also uses indoor humidity, comprehensive comfort, etc. as adjustment targets to meet the user's humidity, comprehensive comfort, etc. needs.
[0081] Reinforcement learning is a machine learning method that enables a networked agent to learn how to make optimal decisions within its environment, thereby achieving self-improvement and optimization. The algorithmic agent analyzes environmental information, then determines appropriate actions to interact with the environment. Finally, the environment returns a corresponding reward to the agent, completing the interaction process. The goal of reinforcement learning algorithm training is to maximize cumulative rewards.
[0082] Specifically, in this embodiment, the environmental information is the input information, namely the operating mode and system status parameters of the air conditioner. The decision action is the output decision parameter, and the air conditioner determines the control action based on the decision parameter to control the operation of the outdoor unit or the air conditioner.
[0083] In this embodiment, a control decision model based on a reinforcement learning algorithm is provided. The control decision model can output decision parameters for adjusting the outdoor unit based on the operating mode and system state parameters. The control decision model is trained using a reinforcement learning algorithm and configured to use the operating mode and system state parameters as input parameters, with at least the rated power utilization rate of the air conditioner's outdoor unit and the indoor temperature as adjustment targets. The control decision model can fully leverage the advantages of the reinforcement learning algorithm in sequential decision-making tasks, deeply exploring the inherent relationship between user-end needs, system decision control, and decision control results. This can achieve energy savings for the outdoor unit while satisfying the comfort requirements of each user, thereby achieving the goal of energy conservation and emission reduction.
[0084] On the other hand, in this embodiment, since the reinforcement learning algorithm accumulates rewards over a longer period, when outputting decision parameters based on the input parameters, the changes in user-end load and control system over a longer period of time will be considered, which can further save the energy consumption of the outdoor unit while improving the stability of the system.
[0085] In some embodiments of the control method of the present invention, the outdoor unit includes at least one compressor, and each compressor is configured to adjust its working state according to at least a high pressure preset value of the outdoor unit or a low pressure preset value of the outdoor unit.
[0086] Compressors can adjust their operating frequency, start and stop times, and other parameters based on preset high or low pressure values. For example, in an outdoor unit with multiple compressors, in cooling mode, the actual low pressure of each compressor is obtained and compared with the preset low pressure value. The compressor's operating frequency can then be appropriately adjusted to ensure that the actual low pressure reaches the preset low pressure value. If all compressors in the active state cannot reach the preset low pressure value by adjusting their operating frequency, additional compressors can be activated.
[0087] In some embodiments of the control method of the present invention, the system state parameters include state parameters and time parameters. The state parameters include outdoor temperature, indoor temperature, frequency of each compressor, power of each compressor, high pressure of each compressor, low pressure of each compressor, heating capacity of the outdoor unit, high pressure preset value, low pressure preset value. The time parameters include time tags of each state parameter.
[0088] In this embodiment, the heat supply of the outdoor unit may be the amount of heating during heating or the amount of cooling during cooling.
[0089] In some embodiments of the control method of the present invention, the decision parameters include a change in a high pressure preset value in the heating mode and a change in a low pressure preset value in the cooling mode.
[0090] The preset value change amount can be a positive value to increase the preset value, or a negative value to decrease the preset value.
[0091] In some embodiments of the control method of the present invention, Figure 2 As shown, the acquisition of the control decision model based on the reinforcement learning algorithm includes:
[0092] S210, building a reinforcement learning model based on the Actor-Critic network;
[0093] S220 , establishing a cyclic interaction between the air conditioner and the reinforcement learning model with a first preset duration as a first cycle, so as to perform online iterative training on the reinforcement learning model and obtain a control decision model.
[0094] The actor-critic network architecture, due to its ability to simultaneously learn both value and policy functions, has become the foundational framework for a number of cutting-edge reinforcement learning algorithms. Reinforcement learning models based on actor-critic networks include Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), and SAC (Soft Actor-Critic), among others, without limitation.
[0095] In this embodiment, the first period can be a single control and adjustment cycle for the air conditioner, for example, 10 minutes. For central air conditioners, multi-split units, and other air conditioners with multiple user-end devices, the control system requires time to reach a stable state after executing a control action. Therefore, the first period can be appropriately set to a longer value to prevent transient state changes in the air conditioner during its response to control commands from interfering with the reinforcement learning model's decision-making.
[0096] During each first cycle, the reinforcement learning model receives environmental state parameters from the air conditioner and outputs control action parameters to the air conditioner. The air conditioner then executes the control action, causing the environmental state parameters to change. The environmental state parameters may include the air conditioner's operating mode and system state parameters, and the control action parameters may include a change in a preset high-pressure pressure value in heating mode and a change in a preset low-pressure pressure value in cooling mode.
[0097] While receiving the environmental state parameters for this cycle, the reinforcement learning model also receives rewards for control actions and accumulates the rewards for each cycle. After the online iterative training is completed, the model with the largest accumulated reward value is selected as the control decision model.
[0098] In some embodiments of the control method of the present invention, Figure 3-4 As shown, before the input parameters are fed into the reinforcement learning model for online training, the following are also included:
[0099] S610, obtaining input parameters of the air conditioner in each first cycle;
[0100] S620: Integrate the system status parameters of each first period.
[0101] The integration process includes:
[0102] S621, interpolating abnormal parameters in the system state parameters to obtain usable system state parameters; abnormal parameters include one or more of missing parameters, parameters not within a preset threshold range, and outlier parameters;
[0103] S622, time-aligning state parameters in each available system state parameter according to each time parameter;
[0104] S623: Resample each state parameter with the first preset time length as a step length.
[0105] In this embodiment, each air conditioner sensor collects operating parameters at its own sampling rate and transmits them to the air conditioner controller or cloud server via wired or wireless means. For example, the data collected by the sensors is uploaded to the device manufacturer's cloud server via a communication module for storage. The format and quality of the raw data collected by the sensors may not necessarily meet the requirements for subsequent environmental interaction and training with the reinforcement learning model. Therefore, some data integration processing is required.
[0106] Specifically, for abnormal data such as missing data, data not within the preset threshold range, and outlier values in the received raw data, interpolation and completion can be adopted. The interpolation and completion methods can be selected from, but not limited to, statistical methods. In the case of continuous missing normal data, where interpolation and completion cannot be performed, data from similar working conditions can be used for completion. If the received data is missing for a long time (a certain threshold can be set, such as 30 minutes), the fault of the corresponding sensor point can be reported for repair processing.
[0107] To normalize the data, the time tags of each parameter also need to be time-aligned. Time alignment is a well-known technique and will not be described in detail here. For example, if data collected and stored at 15:33:38 is time-aligned to data at 15:33:30, and if data collected and stored at 15:33:02 is time-aligned to data at 15:33:00, the data is time-aligned to data at 15:33:00.
[0108] In the prior art, the sensor's acquisition frequency is relatively high, usually once every 30 seconds, which is much less than the first cycle. For this reason, it needs to be resampled. Specifically, the available system state parameters are downsampled. Downsampling is a prior art and will not be described in detail here. Here is only an exemplary explanation: for example, the first cycle is 10 minutes, and the high-pressure pressure parameter acquisition frequency is once every 30 seconds. The high-pressure pressure parameters collected in the first cycle can be averaged and used as the high-pressure pressure parameter value of the first cycle. Alternatively, the high-pressure pressure parameters within 5 minutes before the end of the first cycle can be averaged and used as the high-pressure pressure parameter value of the first cycle.
[0109] Resampling can reduce the amount of data and computing power required. It can also eliminate the transient nature of some sensor data changes, preventing fluctuations from affecting the decision-making of the reinforcement learning model.
[0110] In some embodiments of the control method of the present invention, Figure 5 As shown, the reinforcement learning model based on the Actor-Critic network is constructed, including:
[0111] S211, building an interactive environment based on air conditioners;
[0112] S212, constructing a return pool to store input parameters and decision parameters;
[0113] S213, build an Actor-Critic network agent.
[0114] The interactive environment is used to interact with the Actor-Critic network agent, enabling online iterative training. The return pool is used to store input parameters and decision parameters, which are then called upon when the Actor-Critic network agent interacts with the interactive environment.
[0115] In some embodiments of the control method of the present invention, Figure 6 As shown, the interactive environment based on the air conditioner is constructed, including:
[0116] S711, setting environmental state parameters, which include input parameters;
[0117] S712, setting control action parameters, which include decision parameters;
[0118] S713: Construct a reward and penalty function, where the reward and penalty function is configured to be at least negatively correlated with the rated power utilization of the outdoor unit and negatively correlated with the difference between the indoor temperature and the desired indoor temperature.
[0119] In this embodiment, the environmental state parameters are used to constrain the exploration boundary conditions of the Actor-Critic network agent. Specifically, the environmental state parameters can be expressed as:
[0120] state={Ps i 、Pd i 、Freq i 、W i ,Q,OutTemp,
[0121] OutHum、IndoorTemp、Pd set 、Ps set , Hour…}
[0122] In the above formula, state is the environmental state parameter, i is the number of compressors in the outdoor unit, i = 1, 2, 3..., Ps is the low pressure of the compressor, Pd is the high pressure of the compressor, Freq is the compressor frequency, W is the compressor power, Q is the heating supply, OutTemp is the outdoor temperature, OutHum is the outdoor humidity, IndoorTemp is the weighted average of the indoor temperature, Pd set Ps is the high pressure setting value of the outdoor unit. setis the outdoor unit low pressure setpoint, and Hour is the time value corresponding to the current state of the environmental state parameter, i.e., the time tag. The control action parameters are used to constrain the activity boundary conditions of the Actor-Critic network agent. Specifically, the control action parameters can be expressed as:
[0123]
[0124] In the above formula, action is the control action parameter, is the change in low pressure setting value, The amount of change in the high pressure setting.
[0125] The reward and punishment function is used to reward and punish the Actor-Critic network agent, thereby achieving the regulation goal. Specifically, the reward and punishment function can be expressed as:
[0126] reward=-αJ Pow +βJ com
[0127]
[0128] J com =f com (X1)
[0129] In the above formula, reward is the reward and punishment function, J Pow is the rated power utilization, J com is the comfort index, α and β are the balance coefficients of energy saving and comfort, w is the total power of the compressor, w max is the total rated power of the compressor, f com is the comfort mapping function, X1 is the input parameter of the comfort mapping function, and its input is some parameters that can map indoor comfort, such as indoor temperature, recommended values of related parameters, etc. com Negatively correlated with the difference between the indoor temperature and the expected indoor temperature.
[0130] In some embodiments of the control method of the present invention, the system state parameter further includes outdoor humidity and indoor humidity; and
[0131] The reward and penalty function is further configured to be negatively correlated with the difference between the indoor humidity and the desired indoor humidity.
[0132] In this embodiment, f com The input parameter X1 includes the indoor humidity, and f com Negatively correlated with the difference between indoor humidity and expected humidity at indoor temperature.
[0133] In some embodiments of the control method of the present invention, the reward and penalty function further includes a system state index, specifically,
[0134] reward=-αJ Pow +βJ com -δJ state
[0135]
[0136] In the above formula, J state is the system state index, P ref is the preset threshold of the system state parameter, and δ is the penalty coefficient.
[0137] Exceeding the preset threshold of the system state parameter may trigger unit protection related, therefore, its corresponding action needs to be punished, so that the reward and punishment function is negatively correlated with the difference between the system state parameter and the preset threshold.
[0138] In some embodiments of the control method of the present invention, the constructing and returning the pool comprises:
[0139] Construct an environmental data database. The environmental database is configured to persistently store the data sets generated each time the agent interacts with the air conditioner. The data sets include at least environmental state parameters, control action parameters, and the results of the reward and penalty functions.
[0140] Specifically, the data group in the environmental data database may have the following form:
[0141] [state,action,reward,next_state]
[0142] In the above formula, state is the environmental state parameter of the first cycle, and next_state is the environmental state parameter of the next first cycle.
[0143] In this embodiment, the return pool persistently stores the data sets generated each time the agent interacts with the air conditioner. This allows the data in the return pool to be reused during subsequent training to adjust the reward and penalty functions, train the critic network, and pre-train the actor network. This improves the utilization of the data in the return pool and enhances the efficiency of reinforcement learning training.
[0144] Specifically, the reward and punishment function requires extensive debugging to determine the most appropriate form for different scenarios. If, following existing techniques, the data set returned to the pool is cached data, existing only during a single training session and not persistently stored, then each adjustment to the reward function requires retraining the Actor-Critic network agent from scratch.
[0145] In this embodiment, the data group put back into the pool is persistently stored. Each time the reward and punishment function is adjusted, it is only necessary to recalculate the reward and punishment function value, so that the data group generated by the previous interaction with the environment can be reused.
[0146] In particular, the data set put back into the persistent storage pool can be directly used for training the Critic network. Some valid data sets can also be used for pre-training the Actor initial network.
[0147] In some embodiments of the control method of the present invention, Figure 7 As shown, the construction of the Actor-Critic network agent includes:
[0148] S721, build the Actor initial network and Critic network;
[0149] S722, obtaining pre-training sample data, the pre-training sample data at least including the operating mode of the air conditioner, system state parameters and historical data of control actions;
[0150] S723: Pre-train the Actor initial network according to the pre-training sample data to obtain the Actor network.
[0151] The Actor network is a decision-making network, and the Critic network is an evaluation network. In this embodiment, during reinforcement learning training, the Actor-Critic network agent needs to interact with the air conditioner, that is, to control the operation of the air conditioner by controlling the action parameters.
[0152] Since control action parameters are randomly reset in the early stages of training to avoid being stuck in local values, which can lead to inefficient strategy improvement, this can also affect user comfort and cause the air conditioner to not meet user needs.
[0153] To this end, in this embodiment, when constructing the actor-critic network agent, pre-training sample data is first obtained to pre-train the initial actor network. The pre-trained actor network is then used as the actor-critic network agent for subsequent online interactive training. This allows the actor-critic network agent to make relatively reasonable control actions in the initial stage, thereby reducing the degradation of user experience caused by unreasonable control actions during the initial debugging phase.
[0154] On the other hand, using the pre-trained Actor network as the initial strategy for training the Actor-Critic network agent can also speed up the training of the Actor-Critic network agent to a certain extent, allowing it to converge to a reasonable strategy more quickly.
[0155] In some embodiments of the control method of the present invention, Figure 8 As shown, the steps of the integration process include:
[0156] S801, real-time data input from sensors of the multi-split air conditioning system;
[0157] S802, obtaining a preset threshold range corresponding to the sensor;
[0158] S803, determine whether the data exceeds the preset threshold range; if so, execute S804; if not, execute S805;
[0159] S804, interpolation processing is performed on missing data and data that is not within the preset threshold range; execute S805
[0160] S805, determine whether there is any outlier data; if so, execute S806; if not, execute S807;
[0161] S806, interpolation processing is performed on the data of the outlier value; S807 is executed
[0162] S807, time-aligning the data;
[0163] S808, resampling the data;
[0164] S809, storing in the environmental data database.
[0165] In some embodiments of the control method of the present invention, Figure 9 As shown, the steps for pre-training the Actor initial network include:
[0166] S901, calling and returning the pool data;
[0167] S902, determine whether there is enough data in the back pool; if so, execute S903; if not, execute S904;
[0168] S903, pre-training the Actor initial network using the data put back into the pool;
[0169] S904, pre-training the Actor initial network using existing laboratory data of this type of air conditioner;
[0170] S905: Get the Actor network.
[0171] Figure 11 is a schematic diagram of a machine-readable storage medium 200 according to an embodiment of the present invention. Figure 11As shown, an embodiment of the present invention further provides a machine-readable storage medium 200 on which a machine executable program 201 is stored. When the machine executable program 201 is executed by the processor 132, the air conditioner control method according to any one of the above embodiments or a combination of embodiments is implemented.
[0172] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any machine-readable storage medium 200 for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor 132, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or used in combination with these instruction execution systems, devices or apparatuses.
[0173] For the purposes of the description of this embodiment, the machine-readable storage medium 200 can be any device that can contain, store, communicate, propagate, or transmit a program for use with an instruction execution system, device, or apparatus, or in conjunction with such instruction execution systems, devices, or apparatuses. More specific examples (a non-exhaustive list) of the machine-readable storage medium 200 include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the machine-readable storage medium 200 can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in other suitable ways as necessary, and then storing it in a memory.
[0174] Figure 12 Schematic diagram of an air conditioner 100 according to an embodiment of the present invention, Figure 12 As shown, an embodiment of the present invention further provides an air conditioner 100, which includes a controller 130. The controller 130 includes a memory 131, a processor 132, and a machine executable program 201 stored in the memory 131 and running on the processor 132. When the processor 132 executes the machine executable program 201, the air conditioner control method according to any of the above embodiments is implemented.
[0175] In this embodiment, the air conditioner 100 is equipped with a control decision model based on a reinforcement learning algorithm. This model outputs decision parameters for adjusting the outdoor unit based on the operating mode and system state parameters. This reduces energy consumption in the outdoor unit while ensuring the comfort of each user, thereby achieving energy conservation and emission reduction. Furthermore, because the reinforcement learning algorithm accumulates rewards over a long period of time, the decision parameters it outputs based on the input parameters take into account changes in user load and control system performance over a longer period of time. This further reduces energy consumption in the air conditioner 100 while also improving system stability.
[0176] Specifically, the controller 130 may include a processor 132 adapted to execute stored instructions, and a memory 131 that provides temporary storage space for the operation of the instructions during operation. The processor 132 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 131 may include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.
[0177] The processor 132 can be connected to an I / O interface (input / output interface) suitable for connecting the air conditioner 100 to one or more I / O devices (input / output devices) through a system interconnect (e.g., PCI, PCI-Express, etc.). The I / O devices may include, for example, a keyboard and a pointing device, wherein the pointing device may include a touch pad or a touch screen, etc.
[0178] Processor 132 can also be linked to a display interface suitable for connecting controller 130 to a display device through a system interconnection. Display device can include a display screen as a built-in component of controller 130. Display device can also include a computer monitor, television or projector, etc. that are externally connected to air conditioner 100. In addition, a network interface controller (NIC) can be suitable for connecting controller 130 to a network through a system interconnection. In certain embodiments, NIC can use any suitable interface or protocol (such as Internet Small Computer System Interface, etc.) to transmit data. The network can be a cellular network, a radio network, a wide area network (WAN), a local area network (LAN) or the Internet, etc. Remote devices can be connected to controller 130 through a network.
[0179] In some embodiments of the air conditioner of the present invention, the air conditioner is a multi-split air conditioner. The outdoor unit of the multi-split air conditioner includes a plurality of compressors, and each compressor is configured to adjust its operating state according to at least a preset high pressure value of the outdoor unit or a preset low pressure value of the outdoor unit.
[0180] Through a control decision model based on a reinforcement learning algorithm, the air conditioner can obtain reasonable decision parameters, namely the change in the high-pressure preset value or the change in the low-pressure preset value, to control the working status of each compressor in the outdoor unit. While meeting the comfort of each user end and the stability of the system, it saves the energy consumption of the outdoor unit and achieves the goal of energy conservation and emission reduction.
[0181] The flowchart provided in this embodiment is not intended to indicate that the operations of the method will be performed in any particular order, or that all operations of the method are included in all every case. In addition, the method may include additional operations. Within the scope of the technical ideas provided by the method of this embodiment, additional changes can be made to the above method.
[0182] The present invention has a plurality of exemplary embodiments, but, without departing from the spirit and scope of the present invention, many other variations or modifications that are consistent with the principles of the present invention can be directly determined or derived from the content disclosed in the present invention. Therefore, the scope of the present invention should be understood and deemed to cover all such other variations or modifications.
Claims
1. A method for controlling an air conditioner, characterized in that: include: Obtaining the operating mode and system status parameters of the air conditioner; Obtaining a control decision model based on a reinforcement learning algorithm; the control decision model is configured to: take the operating mode and the system state parameter as input parameters, take at least the rated power utilization rate of the outdoor unit of the air conditioner and the indoor temperature as adjustment targets, and output a decision parameter for adjusting at least the outdoor unit; Obtaining the decision parameter according to the input parameter and the control decision model; A control action is determined according to the decision parameter to control the air conditioner.
2. The control method according to claim 1, characterized in that: The outdoor unit includes at least one compressor, and each compressor is configured to adjust its working state according to at least a high pressure preset value of the outdoor unit or a low pressure preset value of the outdoor unit; and The working mode includes cooling mode and heating mode; The system state parameters include state parameters and time parameters, the state parameters include outdoor temperature, indoor temperature, frequency of each compressor, power of each compressor, high pressure of each compressor, low pressure of each compressor, heating capacity of the outdoor unit, high pressure preset value, low pressure preset value, and the time parameters include time tags of each state parameter; The decision parameters include a change amount of a high pressure preset value in the heating mode and a change amount of a low pressure preset value in the cooling mode.
3. The control method according to claim 2, characterized in that: The acquisition of the control decision model based on the reinforcement learning algorithm includes: Build a reinforcement learning model based on the Actor-Critic network; With a first preset time length as a first cycle, a cyclic interaction is established between the air conditioner and the reinforcement learning model to perform online iterative training on the reinforcement learning model and obtain the control decision model.
4. The control method according to claim 3, characterized in that: Before inputting the input parameters into the reinforcement learning model for online training, the method further includes: acquiring the input parameters of the air conditioner in each of the first cycles; performing integration processing on the system state parameters of each of the first cycles; The integration process includes: Performing interpolation processing on abnormal parameters in the system state parameters to obtain usable system state parameters; the abnormal parameters include one or more of missing parameters, parameters not within a preset threshold range, and outlier numerical parameters; performing time alignment on the state parameters in the available system state parameters according to the time parameters; Resampling is performed on each of the state parameters with the first preset time length as a step length.
5. The control method according to claim 3, characterized in that: The reinforcement learning model based on the Actor-Critic network is constructed, including: Constructing an interactive environment based on the air conditioner; Constructing a return pool to store the input parameters and the decision parameters; Build an Actor-Critic network agent.
6. The control method according to claim 5, characterized in that: The construction is based on the interactive environment of the air conditioner, including: Setting environmental state parameters, wherein the environmental state parameters include the input parameters; Setting control action parameters, wherein the control action parameters include the decision parameters; A reward and penalty function is constructed, where the reward and penalty function is configured to be negatively correlated with at least the rated power utilization of the outdoor unit and negatively correlated with the difference between the indoor temperature and the desired indoor temperature.
7. The control method according to claim 6, characterized in that: The construction is put back into the pool, including: An environmental data database is constructed; the environmental database is configured to persistently store a data group generated each time the intelligent agent interacts with the air conditioner; the data group includes at least the environmental state parameters, the control action parameters and the result of the reward and punishment function.
8. The control method according to claim 6, characterized in that: The construction of the Actor-Critic network agent includes: Build the Actor initial network and Critic network; Acquire pre-training sample data, wherein the pre-training sample data at least includes the operating mode of the air conditioner, the system state parameters, and historical data of the control action; The Actor initial network is pre-trained according to the pre-training sample data to obtain an Actor network.
9. The control method according to claim 7, characterized in that: The system status parameters also include outdoor humidity and indoor humidity; and The reward and penalty function is further configured to be negatively correlated with a difference between the indoor humidity and a desired indoor humidity.
10. The control method according to claim 7, characterized in that: The construction of the reward and punishment function includes: Obtaining a preset threshold value of the system status parameter; The reward and punishment function is configured to be at least negatively correlated with the difference between the system state parameter and the preset threshold.
11. A machine-readable storage medium, characterized in that A machine executable program is stored thereon, and when the machine executable program is executed by a processor, the air conditioner control method according to any one of claims 1 to 10 is implemented.
12. An air conditioner, characterized in that: The invention comprises a controller, wherein the controller comprises a memory, a processor and a machine executable program stored in the memory and running on the processor, and when the processor executes the machine executable program, the control method of the air conditioner according to any one of claims 1 to 10 is implemented.
13. The air conditioner according to claim 12, wherein: The air conditioner is a multi-split unit; the outdoor unit of the multi-split unit includes multiple compressors, and each compressor is configured to adjust its working state according to at least a high-pressure preset value of the outdoor unit or a low-pressure preset value of the outdoor unit.
Citation Information
Patent Citations
Heat pump system with multi-stage compression
CA2597260A1
Indoor fan control method and device
CN106482295A
Multi-split air-conditioning system and pressure detection method
CN108870701A
Robot control method based on offline model pre-training learning DDPG algorithm
CN112668235A
Central air conditioner control method and control system based on reinforcement learning
CN114234381A