Air conditioner and control method and device thereof, storage medium and computer program product
Through reinforcement learning algorithms, optimize the operating parameters of air conditioners and dynamically adjust controllable parameters, the problems of high energy consumption and temperature fluctuations of air conditioners are solved, and energy consumption reduction and temperature stability are achieved.
Patent Information
- Application Number
- CN202510689975.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing air conditioning control methods do not consider multivariate dynamic optimization, resulting in excessive energy consumption or temperature fluctuations, and cannot adapt to environmental changes.
The reinforcement learning algorithm is used to define the state space and action space, and the action value table is generated through the Q-learning algorithm, and the controllable parameters of the air conditioner are dynamically adjusted, such as compressor frequency, outdoor fan speed and expansion valve opening, and combined with the reward function to optimize the control strategy.
It realizes the minimized energy consumption and temperature stability of the air conditioner when meeting user-set parameters, and improves control accuracy and user comfort.
Smart Images

Figure CN120292690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of control, and particularly to an air conditioner and its control method, device, storage medium, and computer program product. Background Art
[0002] Household air conditioners generally adopt variable frequency technology to achieve temperature control by adjusting the compressor frequency, fan speed, and expansion valve opening. Sensors collect real-time data such as indoor and outdoor temperature and humidity, and the controller adjusts the operating parameters according to preset rules, with the goal of maintaining the user-set temperature and reducing energy consumption.
[0003] In related technologies, most variable frequency air conditioners adopt PID control (Proportion-Integral-Differential), but do not consider multi-variable dynamic optimization (such as humidity, outdoor temperature), and the adjustment only depends on static mapping, which easily leads to excessive energy consumption or temperature fluctuations. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the defects of the above-mentioned related technologies, and provide an air conditioner and its control method, device, storage medium, and computer program product to solve the problem that the air conditioner control method in related technologies does not consider dynamic factors and cannot adapt to environmental changes.
[0005] On the one hand, the present invention provides a control method for an air conditioner, including: collecting the current state parameters of the air conditioner; obtaining an action value table that pre-sets the air conditioner to perform different actions in different states; determining the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table; controlling the air conditioner to perform the optimal action.
[0006] Optionally, it further includes: setting the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm.
[0007] Optionally, setting the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm includes: determining the state space and action space of the air conditioner; the state space of the air conditioner includes: a set of different parameter value combinations of two or more state parameters of the air conditioner; the action space of the air conditioner includes: a set of different adjustment range combinations of two or more controllable parameters of the air conditioner; based on the determined state space and action space, calculating the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm.
[0008] Optionally, based on the determined state space and action space, an action value table for the air conditioner to perform different actions in different states is calculated through a reinforcement learning algorithm, including: Initialization step: Initialize the action value table according to the state space and the action space, and set all values of the action value table to 0; Selection step: In each time step, select an action from the action value table according to the current state; Execution step: Execute the selected action and obtain a new state and a reward; The reward is calculated according to a preset reward function; The reward function is set based on the set parameters of the air conditioner; Update step: Update the action value according to the obtained new state and reward according to a preset update rule; Repeat the selection step, the execution step, and the update step until a stop condition is met to obtain the action value table for the air conditioner to perform different actions in different states.
[0009] Optionally, the set parameters include: set temperature and set indoor fan speed; The reward function R is:
[0010] R = -(a * |T_indoor - T_target| + b * |S_indoor_fan - S_indoor_fan_target| + c * P);
[0011] Wherein, T_indoor is the indoor temperature, T_target is the set temperature, S_indoor_fan is the indoor fan speed, S_indoor_fan_target| is the set indoor fan speed, P is the power of the air conditioner, and a, b, and c are the weights of the indoor temperature, indoor fan speed, and power respectively.
[0012] Optionally, record the energy consumption during the actual operation of the air conditioner and / or the temperature difference between the indoor temperature and the set temperature, and verify whether the reduction of the energy consumption of the air conditioner reaches a first preset percentage threshold and / or whether the temperature difference between the indoor temperature and the set temperature meets a preset condition; If the reduction of the energy consumption of the air conditioner does not reach the first preset percentage threshold, increase the weight of the power in the reward function; And / or, if the temperature difference between the indoor temperature and the set temperature meets the preset condition, increase the weight of the indoor temperature in the reward function; Wherein, the preset condition includes: the time ratio of the temperature difference between the indoor temperature and the set temperature being less than a preset temperature value is greater than a second preset percentage threshold.
[0013] Optionally, determining the state space of the air conditioner includes: after discretizing the parameter values of two or more state parameters of the air conditioner according to a preset discretization rule, obtaining a set of different parameter value combinations of the discretized two or more state parameters of the air conditioner, and each parameter value combination of the two or more discretized state parameters is used as a discrete state index in the action value table; determining the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table, including: converting the collected state parameters into discrete state indexes according to the preset discretization rule to match the states in the action value table; and finding the action with the largest action value under the discrete state index in the action value table as the optimal action of the air conditioner in the current state.
[0014] Optionally, it further includes: controlling the indoor blower to operate at a set indoor blower speed.
[0015] Another aspect of the present invention provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any of the foregoing methods are implemented.
[0016] Another aspect of the present invention provides an air conditioner, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of any of the foregoing methods are implemented.
[0017] Another aspect of the present invention provides an air conditioner, including any of the foregoing control devices.
[0018] Another aspect of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the foregoing methods are implemented.
[0019] According to the technical solution of the present invention, the operation parameters of the air conditioner are optimized by using a reinforcement learning algorithm. By collecting state data in real time, the controllable parameters of the air conditioner are dynamically adjusted, and it can adapt to environmental changes, and minimize energy consumption under the condition of meeting the user's set parameters (such as the set indoor temperature and the set indoor blower speed).
[0020] According to the technical solution of the present invention, by defining a state space, an action space, and a reward function, and using Q-learning to generate an optimal control strategy based on a trial-and-error and reward mechanism, the parameters of the air conditioner can be dynamically optimized, the energy consumption of the air conditioner can be reduced, and the temperature stability can be improved.
[0021] According to the technical solution of the present invention, the control accuracy can be improved through a multi-dimensional state space and real-time adjustment; by setting a reward function based on the set parameters, the comfort and energy consumption can be balanced, the user's set goals can be met, and the user comfort can be improved. Description of the Drawings
[0022] The accompanying drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0023] Figure 1 It is a schematic diagram of a method of an embodiment of the control method of the air conditioner provided by the present invention;
[0024] Figure 2 It shows a schematic flowchart of a specific implementation manner of an action value table for setting different actions to be performed in different states of the air conditioner through a reinforcement learning algorithm;
[0025] Figure 3 It shows a flowchart of a specific implementation manner of steps for determining an optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table;
[0026] Figure 4 It shows an overall algorithm development flowchart of the present invention. Specific Embodiments
[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] The present invention provides a control method for an air conditioner.
[0030] Figure 1 It is a schematic diagram of a method of an embodiment of the control method of the air conditioner provided by the present invention.
[0031] As Figure 1 shown, according to an embodiment of the present invention, the control method of the air conditioner at least includes step S110, step S120, step S130, and step S140.
[0032] Step S110, collect the current state parameters of the air conditioner.
[0033] The state parameters may specifically include: the environmental parameters of the environment where the air conditioner is located, as well as the set parameters and operating parameters of the air conditioner. The environmental parameters may specifically include at least one of indoor temperature, outdoor temperature, and indoor humidity; the set parameters may specifically include at least one of set temperature and set indoor fan speed; the operating parameters may specifically include at least one of compressor frequency, indoor fan speed, outdoor fan speed, throttle device opening, and instantaneous power.
[0034] Step S120, obtain the action value table of the air conditioner performing different actions in different states set in advance.
[0035] The action value table of the air conditioner performing different actions in different states, that is, the relationship table between performing different actions and the corresponding values in different states. In a specific embodiment, the action value table of the air conditioner performing different actions in different states is set in advance through a reinforcement learning algorithm (Q-learning algorithm), that is, the Q table.
[0036] Figure 2 shows a schematic flowchart of a specific embodiment of setting the action value table of the air conditioner performing different actions in different states through a reinforcement learning algorithm. As Figure 2 shown, setting the action value table of the air conditioner performing different actions in different states through a reinforcement learning algorithm may specifically include step S1 and step S2.
[0037] Step S1, determine the state space and action space of the air conditioner.
[0038] The state space represents the set of all possible states of the environment. The state space of the air conditioner may specifically include: the set of different parameter value combinations of two or more state parameters of the air conditioner.
[0039] The state parameters may specifically include: the environmental parameters of the environment where the air conditioner is located, as well as the set parameters and operating parameters of the air conditioner. The environmental parameters may specifically include at least one of the indoor temperature T_indoor, the outdoor temperature T_outdoor, and the indoor humidity H_indoor; the set parameters may specifically include at least one of the set temperature T_target and the set indoor fan speed S_indoor_fan_target; the operating parameters may specifically include at least one of the compressor frequency F_compressor, the indoor fan speed S_indoor_fan, the outdoor fan speed S_outdoor_fan, the throttle device opening O_expansion_valve, and the instantaneous power P.
[0040] In a specific embodiment, the state space is composed of different parameter value combinations of the above 10 parameters. These parameters are the state variables that the reinforcement learning agent in the air conditioner system needs to consider, jointly depicting the external conditions of the environment during air conditioner operation (such as indoor and outdoor temperatures, humidity) and the internal working parameters of the system (such as compressor frequency, indoor and outdoor fan speeds, expansion valve opening).
[0041] The state space S can be expressed as:
[0042] S = {T_indoor, T_target, T_outdoor, H_indoor, F_compressor, S_indoor_fan, S_indoor_fan_target, S_outdoor_fan, O_expansion_valve, P}:
[0043] The state space describes all possible situations of the air conditioner and the environment where it is located. Each state is a value combination of these variables (state parameters) at a specific moment, representing the current specific state of the air conditioner system. For example, a state may be "the indoor temperature is 25°C, the target temperature (set temperature) is 22°C, the outdoor temperature is 30°C, the humidity is 50%, and the compressor frequency is 60Hz". The state space S is the set of all these parameter value combinations.
[0044] Preferably, after discretizing the parameter values of two or more state parameters of the air conditioner according to a preset discretization rule, a set of different parameter value combinations of the discretized two or more state parameters of the air conditioner is obtained. Each parameter value combination of the discretized two or more state parameters is used as a discrete state index in the action value table. Specifically, since the parameter values of the state parameters are usually continuous, the values of the state parameters can be discretized. Different state parameters correspond to different discretization rules.
[0045] For example, for the indoor temperature T_indoor, outdoor temperature T_outdoor, and set temperature T_target, it is in steps of 1 °C (e.g., 20 °C, 21 °C, 22 °C, etc.); for the indoor humidity H_indoor, it is in steps of 5% RH (e.g., 40%, 45%, 50%, etc.); for the compressor frequency F_compressor (e.g., 50 Hz, 55 Hz, 60 Hz, etc.), it is in steps of 5 Hz; for the indoor fan speed S_indoor_fan, set indoor fan speed S_indoor_fan_target, and outdoor fan speed S_outdoor_fan, it is in steps of 100 RPM (e.g., 800 RPM, 900 RPM, 1000 RPM, etc.); for the opening O_expansion_valve of the throttling device (e.g., expansion valve), it is in steps of 5% (e.g., 50%, 55%, 60%, etc.).
[0046] In a specific embodiment, after discretizing the parameter values of two or more state parameters of the air conditioner according to a preset discretization rule, each parameter value combination of the two or more discretized state parameters can be represented by a state number as a discrete state index in the action value table.
[0047] Specifically, the discrete state index is a process of converting continuous or complex state parameters of the air conditioner (such as temperature, humidity, etc.) into a finite number of simplified digital tags. Its function is to classify the diverse states in the actual environment into finite and easily manageable state numbers for the system to quickly search and make decisions. Since the state parameters of the air conditioner (such as temperature, wind speed) may originally be continuous values (such as 25.3 °C, wind speed level 3), directly processing all possible values will lead to an excessive amount of calculation. Discretize these values into several intervals. For example, divide the temperature into "low temperature (20 - 22 °C)", "medium temperature (23 - 25 °C)", "high temperature (26 - 28 °C)", and each interval corresponds to a simplified number (such as 1, 2, 3).
[0048] The following uses an example to illustrate how to generate a discrete state index:
[0049] First, segment each parameter. For example:
[0050] Temperature → Discretized into 3 levels (1: low temperature, 2: medium temperature, 3: high temperature);
[0051] Wind speed → Discretized into 2 levels (1: low speed, 2: high speed)
[0052] Then, generate a unique index by combining parameter values. Combine the discrete values of each parameter to form a unique number. For example:
[0053] Temperature = 2 (medium temperature), wind speed = 1 (low speed) → The combined index is "2-1", which can be encoded as the digital index 5.
[0054] The number of states included in the state space of the air conditioner is equal to the product of the number of parameter values of each state parameter. For example, if there are 10 state parameters and each state parameter is discretized into 5 parameter values according to the preset discretization rule, then the number of states included in the state space is 10 5 ones.
[0055] The action space represents the set of all possible actions that can be selected at each decision-making moment. The action space of the air conditioner can specifically include: the set of different adjustment range combinations of two or more controllable parameters of the air conditioner, that is, the action space of the air conditioner is the set of all possible combinations of the adjustment ranges of the two or more controllable parameters. The controllable parameters can specifically include: the compressor frequency F_compressor (for example, in Hz), the outdoor fan speed S_outdoor_fan (for example, in RPM), and the throttle device opening O_expansion_valve (for example, in percentage %).
[0056] The adjustment range of each controllable parameter, that is, the adjustment amount that can be selected when adjusting each controllable parameter. The adjustment amount is the possible amount for adjusting the corresponding controllable parameter once. The adjustment amount of each controllable parameter is discrete, and the set of adjustment ranges of each controllable parameter can be predefined. The set of adjustment ranges of each controllable parameter includes two or more adjustment ranges (i.e., adjustment amounts) that can be selected when adjusting the controllable parameter. For example, the adjustment amount ΔF_compressor of the compressor frequency (that is, the selectable adjustment amount for adjusting the compressor frequency) can include: {-10, -5, 0, +5, +10}, unit: HZ; the adjustment amount ΔS_outdoor_fan of the outdoor fan speed (that is, the selectable adjustment amount for adjusting the outdoor fan speed) can include: {-200, -100, 0, +100, +200}, unit: RPM; the adjustment amount ΔO_expansion_valve of the throttle device opening (for example, the expansion valve opening) (that is, the selectable adjustment amount for adjusting the throttle device opening) can include: {-10, -5, 0, +5, +10}, unit: %. The indoor fan speed is set by the user and is therefore not included in the action space. For an air conditioner system, this discrete adjustment method meets the requirements of actual hardware control.
[0057] In some specific embodiments, the action space A can be represented as:
[0058] A = {ΔF_compressor, ΔS_outdoor_fan, ΔO_expansion_valve}。
[0059] Each action consists of a set of adjustment amounts of various controllable parameters. That is, a set of adjustment amounts of various controllable parameters forms an action combination. For example, for 3 controllable parameters, each controllable parameter's adjustment range includes 5 adjustment amounts, and each action is a triple composed of the adjustment amounts of these 3 parameters. For example, {+5Hz, -100RPM, +5%} is an action, indicating increasing the compressor frequency by 5Hz, decreasing the outdoor fan speed by 100RPM, and increasing the expansion valve opening by 5%. {0Hz, +200RPM, -5%} is another action, indicating keeping the compressor frequency unchanged, increasing the outdoor fan speed by 200RPM, and decreasing the expansion valve opening by 5%.
[0060] The total number of actions included in the action space is equal to the product of the number of selectable adjustment amounts of each controllable parameter. For example, for 3 controllable parameters, each controllable parameter has 5 adjustment amounts. Since there are 5 choices for the adjustment amount of each controllable parameter (for example, ΔF_compressor has 5 values: -10, -5, 0, +5, +10Hz), and the adjustments of these 3 parameters are independent of each other, the total number of actions in the action space is: 5×5×5 = 5 3 = 125, which means the action space contains 125 different action combinations.
[0061] Step S2, based on the determined state space and action space, calculate the action value table of the air conditioner performing different actions in different states through a reinforcement learning algorithm.
[0062] Specifically, at each time step (for example, every minute), according to the current state of the air conditioner system, select an action from the action space to execute. This action will adjust the controllable parameters of the air conditioner, such as the compressor frequency, outdoor fan speed, and expansion valve opening, thereby changing the operating state of the system and ultimately affecting objectives such as comfort and energy consumption. That is, through learning, find out which action combinations can optimize the system performance in different states. That is, through iterative calculation using a reinforcement learning algorithm (Q-learning algorithm), obtain the values of the air conditioner performing different actions in different states in the state space.
[0063] In a specific embodiment, the step of calculating the action value table of the air conditioner performing different actions in different states through a reinforcement learning algorithm based on the determined state space and action space may specifically include an initialization step, a selection step, an execution step, and an update step.
[0064] Initialization step: Initialize the action-value table according to the state space and the action space, and set all values of the action-value table to 0.
[0065] Specifically, create a Q-table (action-value table) according to the state space and the action space, and initialize all values in the table to 0. More specifically, the action-value table (Q-table) is a two-dimensional array, where the columns represent states and the rows represent actions, and each cell Q(s,a) represents the value of performing action a in state s. During initialization, for all s and a, Q(s,a) = 0. The size of the action-value table is equal to the product of the number of states included in the state space and the number of actions included in the action space. That is, the size of the Q-table is equal to the number of states × the number of actions. For example, the number of states included in the aforementioned state space is 10 5 and the number of actions included in the aforementioned action space is 125, then the size of the Q-table is 10 6 × 125.
[0066] Selection step: In each time step, select an action from the action-value table according to the current state.
[0067] Specifically, the ε-greedy strategy can be used to select the action with the highest current Q value with a probability of 1 - ε and randomly select an action with a probability of ε.
[0068] Execution step: Execute the selected action and obtain the new state and reward.
[0069] The reward R can be calculated according to a preset reward function. Specifically, the reward function is set based on the set parameters to balance comfort and the energy consumption of the air conditioner. The set parameters may specifically include the set temperature T_target and the set indoor fan speed S_indoor_fan_target. The set temperature is the target temperature, for example, the temperature set by the user; the set indoor fan speed, for example, the indoor fan speed set by the user.
[0070] In a specific implementation, the reward function R is set as:
[0071] R = -(a * |T_indoor - T_target| + b * |S_indoor_fan - S_indoor_fan_target| + c * P)
[0072] Among them, T_indoor is the indoor temperature, T_target is the set temperature (i.e., the target temperature), S_indoor_fan is the indoor fan speed, S_indoor_fan_target is the set indoor fan speed, and P is the power of the air conditioner. a, b, and c are the weights of the indoor temperature, indoor fan speed, and power respectively. a is the score to be deducted for each difference of the preset temperature value between the indoor temperature and the set temperature, b is the score to be deducted for each difference of the preset speed value between the indoor fan speed and the set fan speed, and c is the score to be deducted for each preset power value of the power.
[0073] For example, k1 = 2 / 1 °C, that is, 2 points are deducted for each 1 °C of the temperature deviation between the indoor temperature and the set temperature, k2 = 1 / 1 RPM, that is, 1 point is deducted for each 1 RPM of the speed deviation between the indoor fan speed and the set fan speed, k3 = 0.01 / 1 W, that is, 0.01 points are deducted for each 1 W of the power. Then:
[0074] R = -(2 * |T_indoor - T_target| + 1 * |S_indoor_fan - S_indoor_fan_target| + 0.01 * P)
[0075] For example, the user sets: T_target = 26 °C, S_indoor_fan_target = 1000 RPM. Current state: T_indoor = 28 °C, S_indoor_fan = 1000 RPM, P = 1100 W. Then:
[0076] R = -(2 * |28 - 26| + 1 * |1000 - 1000| + 0.01 * 1100) = -(4 + 0 + 11) = -15.
[0077] The present invention sets a reward function based on set parameters, can balance comfort and energy consumption, meet the set goals of users, and improve user comfort. The set parameters include the set temperature and the set indoor fan speed, can balance the temperature and the fan speed, meet the set goals of users, and improve user comfort.
[0078] Update step: Update the action value according to the obtained new state and reward according to the preset update rule.
[0079] The preset update rule can specifically be the Q-learning update formula. The Q-learning update formula can be expressed as:
[0080] Q(s,a) ← Q(s,a) + α[R + γmaxQ(s',a') - Q(s,a)]
[0081] Among them, s represents any state in the state space; a represents any action in the action space A of the air conditioner; Q(s, a) represents the value of executing action a in state s, that is, the long-term cumulative reward expectation of executing action a in state s; R represents the immediate reward obtained after executing action a (calculated according to the aforementioned reward function); s′ represents the next state (the new state after executing action a); a' represents the next action selected in the new state s′; α represents the learning rate, which is a fixed value, for example, it can be 0.1; γ represents the discount factor, which is a fixed value, for example, it can be 0.9; maxQ(s', a') represents the value of the optimal action in the new state.
[0082] Repeat and iterate to execute the selection step, the execution step, and the update step until the stop condition is met, and obtain the action value table of the air conditioner executing different actions in different states.
[0083] The stop condition is, for example, the number of iterations and / or the convergence condition. For example, iterate 10 5 times, and the convergence condition is that the change in Q value after 10 5 updates is < 0.01.
[0084] The following is a specific example:
[0085] User setting: T_target = 24°C, S_indoor_fan_target = 800 RPM.
[0086] Initial state:
[0087] T_indoor = 27°C, T_outdoor = 33°C, H_indoor = 55%, P = 1300W.
[0088] F_compressor = 50Hz, S_indoor_fan = 800 RPM, S_outdoor_fan = 700 RPM, O_expansion_valve = 40%.
[0089] Action: {+10Hz, +100 RPM, +5%}.
[0090] New state:
[0091] T_indoor = 25°C, S_indoor_fan = 800 RPM, P = 1400W.
[0092] Reward:
[0093] R = (2 * |25 - 24| + 1 * |800 - 800| + 0.01 * 1400) = (2 + 0 + 14) = -16.
[0094] Q-table update: Adjust the Q-table based on the new state and reward.
[0095] The action value table can be stored in the controller of the air conditioner. For example, the Q-table can be burned into the Flash in the controller.
[0096] Step S130, determine the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table.
[0097] Specifically, according to the collected state parameters, in the action value table for performing different actions in different states of the obtained air conditioner, find the action with the largest action value corresponding to the state parameters, and determine it as the optimal action of the air conditioner in the current state.
[0098] Figure 3 Shows a flowchart of a specific implementation of the step of determining the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table. As Figure 3 shown, in a specific implementation, step S130 may include the following steps S131 and S132.
[0099] Step S131, convert the collected state parameters into discrete state indices according to the preset discretization rule to match the states in the action value table.
[0100] Step S132, according to the converted discrete state index, find the action with the largest action value of the air conditioner in the current state in the action value table as the optimal action of the air conditioner in the current state.
[0101] Since the set of different parameter value combinations of two or more state parameters of the air conditioner included in the state space is obtained by discretizing the parameter values of two or more state parameters of the air conditioner according to the preset discretization rule, and each parameter value combination of the two or more state parameters after discretization is used as a discrete state index in the action value table, therefore, the collected state parameters can be discretized according to the preset discretization rule and converted into discrete state indices first, and then find the action with the largest action value corresponding to the converted discrete state index in the action value table as the optimal action of the air conditioner in the current state.
[0102] Step S140, control the air conditioner to execute the optimal action.
[0103] That is, control the adjustment amounts of the controllable parameters corresponding to the optimal actions of the air conditioner. For example, adjust the compressor frequency F_compressor, the outdoor fan speed S_outdoor_fan, and the opening degree O_expansion_valve of the throttling device, thereby reducing energy consumption.
[0104] Preferably, the controllable parameters need to be controlled within a preset range. For example, the range of the compressor frequency F_compressor can be, for example: 20 - 120 Hz, the range of the outdoor fan speed S_outdoor_fan can be, for example: 500 - 1200 RPM, and the range of the opening degree O_expansion_valve of the throttling device can be: 0 - 100%.
[0105] Preferably, the method may further include: controlling the indoor fan to operate at the set indoor fan speed. The indoor fan speed is set by the user and is not included in the action space. Therefore, the indoor fan speed can be forced to be set to the target speed (S_indoor_fan_target) set by the user, that is, set the indoor fan speed to ensure user priority and meet the user's comfort requirements.
[0106] Optionally, the method may further include: during the operation of the air conditioner, if the temperature deviation of the indoor temperature T_indoor from the set temperature T_target exceeds a preset temperature value, enter the full power mode. For example, if T_indoor deviates from T_target by more than 5°C, enter the full power mode.
[0107] Preferably, the method may further include: recording the energy consumption during the actual operation of the air conditioner and / or the temperature difference between the indoor temperature and the set temperature, and verifying whether the reduction of the energy consumption of the air conditioner reaches a first preset percentage threshold and / or whether the temperature difference between the indoor temperature and the set temperature meets a preset condition; if the reduction of the energy consumption of the air conditioner does not reach the first preset percentage threshold, increase the weight of the power in the reward function; and / or, if the temperature difference between the indoor temperature and the set temperature meets the preset condition, increase the weight of the indoor temperature in the reward function. The preset condition may specifically be that the proportion of the time when the temperature difference between the indoor temperature and the set temperature is less than the preset temperature value is greater than a second preset percentage threshold.
[0108] For example, run for 7 days and record the energy consumption and comfort level when the air conditioner is running. The target for energy consumption reduction is 15% - 20% (compared with the energy consumption under the same set temperature T_target and set indoor fan speed S_indoor_fan_target). The preset condition for the temperature difference between the indoor temperature and the set temperature is that the proportion of the temperature difference between the indoor temperature and the set temperature < 1°C is greater than 90%. If the energy consumption does not meet the standard, increase the power weight (for example, c increases from 0.01 to 0.015). If the temperature difference between the indoor temperature and the set temperature is too large, increase the temperature weight (for example, a increases from 2 to 3).
[0109] To clearly illustrate the technical solution of the present invention, the execution process of the control method for the air conditioner provided by the present invention will be described below with a specific embodiment.
[0110] According to a specific embodiment of the present invention, the method includes the following steps:
[0111] Collect status: Read 10 variables of the current status, including indoor temperature (T_indoor), target temperature (T_target), outdoor temperature (T_outdoor), indoor humidity (H_indoor), compressor frequency (F_compressor), indoor fan speed (S_indoor_fan), target fan speed (S_indoor_fan_target), outdoor fan speed (S_outdoor_fan), expansion valve opening (O_valve), and instantaneous power (P), and store them as a status array.
[0112] Status discretization: Convert the continuous status values into discrete indices (s_idx) to match the status in the Q table.
[0113] Select action: According to the Q table, select the action (action) with the largest Q value under the current status index (s_idx).
[0114] Execute action: Execute the selected action and adjust the compressor frequency, outdoor fan speed, and expansion valve opening.
[0115] Set fan speed: Force the indoor fan speed to be set to the user-specified target speed (S_indoor_fan_target) to ensure user settings take precedence.
[0116] Wait for 60 seconds (one cycle per minute), and then return to continue the next cycle. The user can update the target temperature (T_target) and target fan speed (S_indoor_fan_target) at any time through the remote control or serial port, and the program responds to these changes in real time.
[0117] The above control logic can also be referred to as follows:
[0118]
[0119] Figure 4 The overall algorithm development process of the present invention can be referred to Figure 4 As shown:
[0120] (1) Install hardware: including sensors such as indoor temperature, humidity sensor, outdoor temperature sensor, power meter, and controller (with appropriate chip specifications, such as 64KB RAM, 256KB Flash).
[0121] Actuators: variable frequency compressor, speed-controlled outdoor fan, electronic expansion valve.
[0122] (2) Data collection: The status parameters of the air conditioner are collected through the above sensors.
[0123] (3) Training model: using reinforcement learning algorithm to train the Q table;
[0124] (4) Burning Q table: The Q table can be burned into the Flash in the controller.
[0125] (5) Real-time control: Collect real-time state parameters, select the action with the largest Q value in the Q table based on the real-time state parameters, and execute it.
[0126] (6) Verification effect: Verify whether the energy consumption is reduced and whether the temperature difference between the indoor temperature and the set temperature meets the requirements.
[0127] (7) Optimization and adjustment: If the energy consumption does not meet the standard, increase the power weight in the reward function; if the temperature difference between the indoor temperature and the set temperature is too large, increase the temperature weight in the reward function.
[0128] The technical solution of the present invention optimizes the compressor frequency, outdoor fan speed and expansion valve opening under the user-set T_target and S_indoor_fan_target through Q-learning, achieving a 15-20% energy consumption reduction. It is operable from data collection to deployment and is suitable for household air conditioning systems that support variable frequency control.
[0129] The present invention also provides a storage medium corresponding to the air conditioner control method, on which a computer program is stored, and when the program is executed by a processor, the steps of any of the above methods are implemented.
[0130] The present invention also provides an air conditioner corresponding to the control method of the air conditioner, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the foregoing methods are implemented.
[0131] The present invention also provides a computer program product corresponding to the control method of the air conditioner, including a computer program. When the computer program is executed by a processor, the steps of any of the foregoing methods are implemented.
[0132] Accordingly, the solution provided by the present invention uses a reinforcement learning algorithm to optimize the operating parameters of the air conditioner. By collecting status data in real time, the controllable parameters of the air conditioner, such as the compressor frequency, the outdoor fan speed, and the expansion valve opening, are dynamically adjusted to minimize energy consumption while meeting the user-set parameters (such as the indoor set temperature and the set indoor fan speed).
[0133] The solution provided by the present invention defines a state space, an action space, and a reward function, and uses Q-learning to generate an optimal control strategy based on a trial-and-error and reward mechanism, which can dynamically optimize the parameters of the air conditioner, reduce the energy consumption of the air conditioner, and improve the temperature stability.
[0134] The solution provided by the present invention can improve the control accuracy through a multi-dimensional state space and real-time adjustment; setting a reward function based on set parameters can balance comfort and energy consumption, meet the user's set goals, and improve user comfort.
[0135] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of the present invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination of these. In addition, each functional unit can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0136] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0137] The units described as separate components may or may not be physically separated. The components serving as control devices may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0139] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A control method for an air conditioner, characterized in that, Including: Collecting the current state parameters of the air conditioner; Obtaining the action value table that the air conditioner performs different actions in different states, which is preset; Determining the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table; Controlling the air conditioner to perform the optimal action.
2. The method according to claim 1, wherein It also includes: Setting the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm.
3. The method according to claim 2, wherein Setting the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm, including: Determining the state space and action space of the air conditioner; the state space of the air conditioner includes a set of different parameter value combinations of more than two state parameters of the air conditioner; the action space of the air conditioner includes a set of different adjustment range combinations of more than two controllable parameters of the air conditioner; Based on the determined state space and action space, calculating the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm.
4. The method according to claim 3, wherein Based on the determined state space and action space, calculating the action value table that the air conditioner performs different actions in different states through a reinforcement learning algorithm, including: Initialization step: Initializing the action value table according to the state space and the action space, and setting all values of the action value table to 0; Selection step: Selecting an action from the action value table according to the current state in each time step; Execution step: Executing the selected action and obtaining a new state and a reward; the reward is calculated according to a preset reward function; the reward function is set based on the set parameters of the air conditioner; Update step: Updating the action value according to the obtained new state and reward according to a preset update rule; Repeatedly iteratively execute the selection step, the execution step and the update step until a stop condition is met, and obtain the action value table that the air conditioner performs different actions in different states.
5. The method according to claim 4, wherein The set parameters include: set temperature and set indoor fan speed; the reward function R is: R = -(a * |T_indoor - T_target| + b * |S_indoor_fan - S_indoor_fan_target| + c * P) Wherein, T_indoor is the indoor temperature, T_target is the set temperature, S_indoor_fan is the indoor fan speed, S_indoor_fan_target| is the set indoor fan speed, P is the power of the air conditioner, and a, b and c are the weights of the indoor temperature, the indoor fan speed and the power respectively.
6. The method according to claim 5, wherein Recording the energy consumption and / or the temperature difference between the indoor temperature and the set temperature during the actual operation of the air conditioner, and verifying whether the reduction of the energy consumption of the air conditioner reaches a first preset percentage threshold and / or whether the temperature difference between the indoor temperature and the set temperature meets a preset condition; If the reduction of the energy consumption of the air conditioner does not reach the first preset percentage threshold, increase the weight of the power in the reward function; and / or, if the temperature difference between the indoor temperature and the set temperature meets the preset conditions, increase the weight of the indoor temperature in the reward function; Wherein, the preset conditions include: the time ratio that the temperature difference between the indoor temperature and the set temperature is less than the preset temperature value is greater than the second preset percentage threshold.
7. The method according to any one of claims 1-6, wherein Determining the state space of the air conditioner includes: after discretizing the parameter values of two or more state parameters of the air conditioner according to the preset discretization rules, obtaining a set of different parameter value combinations of the discretized two or more state parameters of the air conditioner, and each parameter value combination of the two or more state parameters after discretization is used as a discrete state index in the action value table; Determining the optimal action of the air conditioner in the current state according to the collected state parameters and the obtained action value table includes: Converting the collected state parameters into discrete state indexes according to the preset discretization rules to match the states in the action value table; Searching for the action with the largest action value under the discrete state index in the action value table as the optimal action of the air conditioner in the current state.
8. The method according to any one of claims 1 to 6, characterized in that, It further includes: Controlling the indoor fan to operate at the set indoor fan speed.
9. A storage medium, characterized in that, It stores a computer program, and when the program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
10. An air conditioner, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and operable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1-8.
11. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Cited By
Method and device for optimizing heat exchange efficiency of air conditioner and air conditioner
CN120991417A