Multi-split air conditioner and control method thereof
By describing the control of multiple online air conditioners as a framework for reinforcement learning, and adopting greedy strategies and exploration strategies, the problem that traditional control methods fail to consider the coupling relationship of the actuator and find the optimal control action is solved, and the energy saving and user comfort of multiple online air conditioners are achieved.
Patent Information
- Application Number
- CN202510257604.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-13
AI Technical Summary
The traditional multi-online air conditioning control method fails to effectively consider the coupling relationship between actuators, resulting in the inability to find the optimal control action in different user scenarios, and it is difficult to take into account both energy saving and user comfort.
The reinforcement learning method is adopted to describe the control of multiple online air conditioners as a framework for reinforcement learning. By identifying the current operating status, selecting target control actions and executing them, calculating energy efficiency ratios, and switching based on greedy strategies and exploration strategies to find the optimal control actions.
It realizes the optimal control action of finding multiple online air conditioners in different user scenarios, taking into account the energy saving and user comfort of multiple online air conditioners.
Smart Images

Figure CN119983494A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of refrigeration, and in particular to a multi-connected air conditioner and a control method thereof. Background Art
[0002] The control of multi-split air conditioners requires cooperation between the internal and external units, with many control variables and complex processes.
[0003] Traditional multi-split air conditioners adjust indoor temperature through PID control strategy, but this control method is only suitable for the case of a single control variable. Therefore, the control of each actuator of the multi-split air conditioner based on the PID control strategy is independent of each other, and the coupling relationship between the actuators is not considered.
[0004] In addition, there is a complex nonlinear relationship between the operating state of the multi-split air conditioner and the actions of the internal and external actuators, which makes it difficult to model them.
[0005] The reinforcement learning method with self-learning and model-free characteristics can find the optimal action for each state based on this mechanism of continuous exploration and interaction by identifying the object state, outputting actions, and calculating the reward value according to the feedback state of the object. It is suitable for the control process of multi-split air conditioners with high complexity.
[0006] Therefore, how to design a multi-split air conditioner and its control method to find the optimal control action of the multi-split air conditioner in different user scenarios is a technical problem that needs to be solved urgently in the industry. Summary of the invention
[0007] In view of the problem in the prior art that a multi-split air conditioner under a traditional PID control strategy does not consider the coupling relationship between actuators and cannot adapt to different user scenarios to find the optimal control action, the present invention proposes a multi-split air conditioner and a control method thereof.
[0008] The technical solution of the present invention is to propose a control method for a multi-split air conditioner, comprising:
[0009] Step S1: identifying the current operating state of the multi-split air conditioner, and acquiring all control actions corresponding to the current operating state according to a preset corresponding relationship;
[0010] Step S2: Based on the first preset strategy, a target control action is selected from all control actions corresponding to the current operating state and executed, and the energy efficiency ratio of the multi-split air conditioner after executing the target control action is calculated in turn, and the target control action corresponding to the highest energy efficiency ratio of the multi-split air conditioner is obtained and executed.
[0011] Furthermore, the first preset strategy includes a greedy strategy and an exploration strategy;
[0012] The greedy strategy is used to select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the executed target control actions, and the exploration strategy is used to select the target control action that has not been executed from all the control actions;
[0013] The greedy strategy and the exploration strategy are switched according to the size relationship between the generated random number and the set parameter.
[0014] Furthermore, the step S2 comprises:
[0015] Step S21: Generate a random number and compare it with the set initial parameter. When the random number is smaller than the initial parameter, execute the exploration strategy and enter step S22. When the random number is larger than the initial parameter, execute the greedy strategy and enter step S23.
[0016] Step S22: Calculate and save the energy efficiency ratio of the multi-split air conditioner after executing the selected target control action, and update the initial parameters at a preset ratio and then return to step S21;
[0017] Step S23: Select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the target control actions that have been executed, compare the corresponding energy efficiency ratio of the multi-split air conditioner with the saved energy efficiency ratio of the multi-split air conditioner, save the larger value in the comparison result, and update the initial parameters at a preset ratio before returning to step S21.
[0018] Furthermore, the step S21 further includes:
[0019] Step S211: After executing the selected target control action, detecting the current indoor temperature;
[0020] Step S212: Determine whether the current indoor temperature exceeds the preset temperature range. If so, control the operating state of the multi-split air conditioner with the second preset strategy. If not, execute the greedy strategy or exploration strategy based on the size relationship between the generated random number and the set parameter.
[0021] Furthermore, the second preset strategy is a PID control strategy.
[0022] Furthermore, the expression model of the current running state is:
[0023]
[0024] in, is the observed state of indoor unit i at time t, N is the number of indoor units, is the current indoor temperature of indoor unit i, is the indoor temperature of indoor unit i at the last moment, is the current set temperature of indoor unit i, θ i is the current indoor humidity of indoor unit i, is the current indoor unit capacity of indoor unit i, FL i is the current windshield of indoor unit i, T out is the current outdoor temperature, P o It is the current operating power of the outdoor unit.
[0025] Furthermore, the expression model of the control action is:
[0026] a t ={EXV 1 ,EXV 2 ,L,EXV N ,f com ,f fan}
[0027] Among them, EXV i is the expansion valve opening of unit i at time t, f com is the compressor frequency of the outdoor unit, f fan is the fan frequency.
[0028] Furthermore, the calculation model of the energy efficiency ratio of the multi-split air conditioner is:
[0029]
[0030] in, is the current internal unit capacity of internal unit i, P o is the current operating power of the outdoor unit, and N is the number of indoor units.
[0031] Furthermore, the initial parameter is set to 0.9, the random number is selected from the interval between 0 and 1, and the preset ratio is 0.99.
[0032] The present invention also proposes a multi-split air conditioner, which has one outdoor unit, three indoor units, and an actuator, and the actuator is used to execute the control method of the multi-split air conditioner.
[0033] Compared with the prior art, the present invention has at least the following beneficial effects:
[0034] The control method of the multi-split air conditioner proposed in the present invention describes the control of the multi-split air conditioner as a reinforcement learning framework. Both its state space and action space take into account multiple states. At the same time, the energy efficiency ratio of the multi-split air conditioner is used as the final reward value. It can find the optimal control action of the multi-split air conditioner in different user scenarios, which not only takes into account the energy saving of the multi-split air conditioner, but also ensures the comfort of the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 This is a schematic diagram of the structure of the multi-split air conditioner proposed by the present invention;
[0037] Figure 2 A control flow chart of the multi-split air conditioner proposed by the present invention;
[0038] Figure 3 It is the relationship Q table corresponding to the state space and the action space in the present invention;
[0039] Figure 4 This is a table corresponding to the reward values of all control actions in running state s2. DETAILED DESCRIPTION
[0040] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0041] Thus, a feature indicated in this specification will be used to illustrate one of the features of an embodiment of the present invention, rather than implying that each embodiment of the present invention must have the described feature. In addition, it should be noted that this specification describes many features. Although some features can be combined together to illustrate possible system designs, these features can also be used in other combinations that are not explicitly described. Thus, unless otherwise stated, the described combinations are not intended to be limiting.
[0042] The principle and structure of the present invention are described in detail below with reference to the accompanying drawings and embodiments.
[0043] At present, the traditional control method of multi-split air conditioner mostly uses PID control strategy, but this control strategy is only suitable for the case of a single control variable. However, multi-split air conditioner often has multiple control variables. Therefore, the control of each actuator of the multi-split air conditioner under the PID control strategy is independent of each other (such as one actuator is used to control the compressor frequency, and another actuator is used to control the fan frequency), and it does not consider the coupling relationship between the actuators. However, for a multi-split air conditioner, the energy efficiency ratio of the whole machine often needs to take into account multiple control variables. For a single control variable, the best energy efficiency ratio can be achieved under a certain control parameter, but it will affect the energy efficiency ratio of other control variables on the multi-split air conditioner, which makes the multi-split air conditioner fail to achieve the best energy efficiency ratio.
[0044] Therefore, when controlling a multi-split air conditioner, it is necessary to consider multiple control variables in order to obtain an accurate optimal energy efficiency ratio. It should be noted that the optimal energy efficiency ratio will also change under different working environments. At present, the reinforcement learning method with self-learning and model-free characteristics can well adapt to the control under multiple control variables. Therefore, the design idea of the present invention is to describe the control of the multi-split air conditioner as a reinforcement learning framework, so as to use the reinforcement learning method to achieve the optimal control of the multi-split air conditioner under multiple control variables.
[0045] Based on the above design ideas, the present invention constructs the operating states of the outdoor unit and indoor unit of the multi-split air conditioner into the state space of reinforcement learning, and constructs the adjustment actions of each actuator, such as the compressor frequency of the outdoor unit, the fan frequency, and the expansion valve opening of each indoor unit, into the action space of reinforcement learning, and constructs the corresponding reward function at the same time. Since the present invention considers the energy efficiency ratio of the multi-split air conditioner, the reward function is actually a calculation function of the energy efficiency ratio of the multi-split air conditioner, and the reward value is also the energy efficiency ratio of the multi-split air conditioner.
[0046] Here, the above-mentioned state space can be understood as including the operating states of multiple multi-split air conditioners, and the action space can be understood as including the control actions of multiple multi-split air conditioners. In order to realize the control of the multi-split air conditioner, it is necessary to first associate the above-mentioned state space and action space, and establish a corresponding relationship Q table, which is the corresponding relationship described in the previous article.
[0047] See also Figure 3, the relationship Q table here is the correspondence between multiple operating states and multiple control actions. It has n operating states and n control actions. Multiple control actions can be executed in each operating state. The control idea of the present invention is to determine the current operating state of the multi-split air conditioner, obtain all the control actions in the operating state from its relationship Q table, and screen them one by one to obtain the control action with the highest energy efficiency ratio, so as to achieve the best control action of the multi-split air conditioner in the current operating state. It should be pointed out that the above operating states and control actions do not refer to a single control variable, but a collection of multiple control variables. Therefore, when controlling the multi-split air conditioner, the present invention can take into account the coupling between multiple control variables.
[0048] Based on the above ideas, the control method of the multi-split air conditioner proposed in the present invention includes the following steps:
[0049] Step S1: Identify the current operating state of the multi-split air conditioner, and obtain all control actions corresponding to the current operating state according to a preset corresponding relationship;
[0050] Step S2: Based on the first preset strategy, select and execute the target control action from all control actions corresponding to the current operating state, calculate the energy efficiency ratio of the multi-split air conditioner after executing the target control action in turn, obtain the target control action corresponding to the highest energy efficiency ratio of the multi-split air conditioner and execute it.
[0051] Here, the corresponding relationship in step S1 is the relationship Q table mentioned above. It can be seen from the above relationship Q table that the multi-split air conditioner has n operating states and n control actions, which are respectively recorded as s1, s2, ..., s n , and a1, a2, ... a n After identifying the current operating state of the multi-split air conditioner, the present invention can obtain multiple control actions that can be executed. For example, if the current operating state of the multi-split air conditioner is s2, it can execute a1, a2, ... a n And other control actions.
[0052] The corresponding R in the relationship Q table 21 , R 22 , ……R 2n is the reward value obtained after executing the corresponding control action under the current operating state s2, that is, the energy efficiency ratio of the multi-split air conditioner, which corresponds to the energy efficiency ratio of the multi-split air conditioner obtained after executing the target control action in the above step S2. At the same time, the largest energy efficiency ratio is selected from the energy efficiency ratios of all the multi-split air conditioners, and the corresponding control action is the optimal control action of the multi-split air conditioner under the current operating state;
[0053] See also Figure 4 , which is the control action corresponding to the current running state s2 in one embodiment of the present invention, from Figure 4 It can be seen that the current running state s2 corresponds to four control actions, namely a1, a2, a3, and a4. Under the control action a1, the specific actions include: the compressor frequency is 58Hz (here refers to the compressor of the outdoor unit of the multi-split air conditioner), the outdoor fan frequency is 42Hz (that is, the fan frequency in the previous text), the valve opening of the indoor unit 1 is 400, and the valve opening of the indoor unit 2 is 460 (in this embodiment, the multi-split air conditioner has only two indoor units, and the valve opening here refers to the expansion valve opening). At this time, the reward value R 21 is 3.63 (that is, the energy efficiency ratio of the multi-split air conditioner is 3.63);
[0054] Under control action a2, the specific actions include: the compressor frequency is 58 Hz, the external fan frequency is 48 Hz, the valve opening of the indoor unit 1 is 420, and the valve opening of the indoor unit 2 is 460. The reward value at this time is unknown (all control actions in the present invention need to be explored and executed based on the first preset strategy. The unknown here is because the control action has not been executed yet);
[0055] Under control action a3, the specific actions include: compressor frequency is 58Hz, external fan frequency is 50Hz, valve opening of indoor unit 1 is 420, valve opening of indoor unit 2 is 480. The reward value R at this time is 23 is 3.60;
[0056] Under control action a4, the specific actions include: compressor frequency is 60Hz, external fan frequency is 48Hz, valve opening of indoor unit 1 is 460, valve opening of indoor unit 2 is 480. The reward value R at this time is 24 is 3.40;
[0057] From the known reward values, when executing control action a1, the calculated reward value is the largest, that is, the energy efficiency ratio of the multi-split air conditioner is the highest (it should be pointed out that since control action a2 has not been explored at this time, if the reward value obtained after executing control action a2 is larger, then in the current operating state, executing control action a2 can make the energy efficiency ratio of the multi-split air conditioner the highest);
[0058] It should be pointed out here that the above-mentioned execution of the corresponding control action is not directly executed. As mentioned above, the present invention is provided with a corresponding reward function. Therefore, after exploring the corresponding control action, the reward function can be directly used to obtain the corresponding reward value without directly executing the control action.
[0059] It can be seen from the above-mentioned overall control scheme that the present invention can obtain the control action with the highest energy efficiency ratio from all the control actions that can be executed according to the current operating state of the multi-split air conditioner, and the control action has multiple control variables, which can ensure that the control action obtained at this time is the optimal control action of the current multi-split air conditioner, which not only takes into account the energy saving of the multi-split air conditioner, but also ensures the comfort of the user.
[0060] That is, the present invention adopts the control method of the multi-split air conditioner, which has at least the following beneficial effects:
[0061] The control method of the multi-split air conditioner proposed in the present invention describes the control of the multi-split air conditioner as a reinforcement learning framework. Both its state space and action space take into account multiple states. At the same time, the energy efficiency ratio of the multi-split air conditioner is used as the final reward value. It can find the optimal control action of the multi-split air conditioner in different user scenarios, which not only takes into account the energy saving of the multi-split air conditioner, but also ensures the comfort of the user.
[0062] It can be seen from the above scheme that when the present invention executes the above control method, the most critical step is how to select the target control action and obtain the target control action corresponding to the highest energy efficiency ratio. Therefore, the first preset strategy for implementing the above action in the present invention is described below:
[0063] The essence of the first preset strategy in the present invention is an ε-greedy strategy, which includes two parts, namely a greedy strategy and an exploration strategy;
[0064] Here, the greedy strategy is used to select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the executed target control actions, and the exploration strategy is used to select the target control action that has not been executed from all the control actions;
[0065] The greedy strategy and exploration strategy are switched according to the size relationship between the generated random number and the set parameters.
[0066] Based on the ε-greedy strategy, the above step S2 includes the following steps:
[0067] Step S21: Generate a random number and compare it with the set initial parameter. When the random number is less than the initial parameter, execute the exploration strategy and enter step S22. When the random number is greater than the initial parameter, execute the greedy strategy and enter step S23.
[0068] Step S22: Calculate and save the energy efficiency ratio of the multi-split air conditioner after executing the selected target control action, and update the initial parameters with a preset ratio and then return to step S21;
[0069] Step S23: Select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the target control actions that have been executed, compare the corresponding energy efficiency ratio of the multi-split air conditioner with the saved energy efficiency ratio of the multi-split air conditioner, save the larger value in the comparison result, and update the initial parameters at a preset ratio before returning to step S21.
[0070] In the present invention, the above-mentioned initial parameter is set to a larger value, which can ensure that when generating random numbers, the random numbers will be smaller than the initial parameter. At the same time, the above-mentioned preset ratio is set to a ratio less than 1. In this way, after each execution of the ε-greedy strategy, when returning to step S21, the initial parameter is reduced, so that the random number can be randomly assigned to a value larger than the initial parameter with a greater probability.
[0071] The above steps are described in detail below. In the early stage of executing the above ε-greedy strategy, since the value of the initial parameter setting is relatively large, the random number generated will basically be smaller than the initial parameter. According to the above logic, the exploration strategy will be executed at this time. Under the exploration strategy, the present invention will select an unexecuted control action from all control actions (it can also be understood as an unselected control action). With the repeated execution of the exploration strategy, the initial parameter will become smaller and smaller, resulting in an increasing probability that the random number is larger than the initial parameter. At this time, the greedy strategy will begin to be executed;
[0072] Under the greedy strategy, the present invention will select the control actions that have been executed (which can also be understood as the selected control actions), and obtain the maximum reward value (that is, the energy efficiency ratio of the multi-split air conditioner). At the same time, it will compare with the previously recorded reward value each time it is executed, and save the maximum value. If it is the first time to execute the greedy strategy, the current reward value will be saved directly. Each subsequent execution of the greedy strategy will be compared with the previously saved reward value, and the larger value will be saved;
[0073] Here, each time the greedy strategy is executed, it will be compared with the reward value saved last time, and the larger value will be saved at the same time, which is used to update the maximum reward value in real time. For example, the reward value saved when the greedy strategy is executed for the first time is 3.6 (the value saved by each greedy strategy is the maximum reward value corresponding to the control action that has been executed), but the exploration strategy will still be executed in the subsequent process, and a larger reward value may appear at this time, such as 3.7, but the greedy strategy will not be executed at the same time. Therefore, when the greedy strategy is executed for the second time, the maximum reward value at this time needs to be compared with the reward value saved last time (that is, the maximum reward value when the exploration strategy is not executed), and then the maximum value is saved, so that the maximum reward value can be updated in real time, that is, the maximum value of the energy efficiency ratio of the multi-split air conditioner.
[0074] Based on the above-mentioned ε-greedy strategy, in the early stage of the strategy, the present invention will basically execute the exploration strategy all the time, so as to explore all the control actions. In the later stage of the strategy, since the exploration strategy has been executed many times, the control actions have basically been explored, and since the initial parameters at this time have become smaller, the greedy strategy will basically be executed all the time, and the reward value will be continuously updated until the maximum reward value is obtained (because the execution of the greedy strategy this time will be compared with the reward value saved by the previous execution of the greedy strategy, and the larger value will be saved, so the final reward value must be the maximum reward value);
[0075] That is, the ε-greedy strategy encourages exploration strategy in the early stage and greedy strategy in the later stage.
[0076] Among them, the above-mentioned preset ratio can be set according to actual needs. If in a certain embodiment, there are many control actions that can be executed in the current operating state, the preset ratio can be set to a value close to 1 but less than 1, so that the initial parameter can be reduced more slowly, thereby executing more exploration strategies in the early stage; similarly, if there are few control actions that can be executed in the current operating state, the preset ratio can be set to a smaller value, thereby increasing the speed of reducing the initial parameter and starting the greedy strategy in advance.
[0077] In a preferred embodiment of the present invention, the above initial parameter is set to 0.9, the random number is selected from the interval of 0 to 1, and the above preset ratio is 0.99.
[0078] Furthermore, under certain specific working conditions, the control method of the multi-split air conditioner proposed in the present invention may cause some control variables to exceed the threshold in order to achieve the highest energy efficiency ratio. In this case, a PID control strategy needs to be adopted for adjustment. To address the above problem, in the control method proposed in the present invention, step S21 further includes:
[0079] Step S211: After executing the selected target control action, detecting the current indoor temperature;
[0080] Step S212: Determine whether the current indoor temperature exceeds the preset temperature range. If so, control the operating state of the multi-split air conditioner with the second preset strategy. If not, execute the greedy strategy or the exploration strategy according to the size relationship between the generated random number and the set parameter.
[0081] Here, the above preset temperature range is set as: Here The average set temperature for the user.
[0082] See also Figure 2 , which is an overall control flow chart of the control method of the multi-split air conditioner in the present invention, which includes the following steps:
[0083] Identify the current state S t ; This step is also the step S1 in the previous text of identifying the current operating status of the multi-split air conditioner;
[0084] Select an action in the Q table based on the ε-greedy strategy and execute it; this step is also the step S2 above, which selects and executes the target control action from all the control actions corresponding to the current running state based on the first preset strategy. Here, the ε-greedy strategy is also the first preset strategy, and the Q table is also the corresponding relationship in the previous text, which expresses all the control actions corresponding to the current running state. Selecting an action and executing it is also selecting and executing the target control action;
[0085] Determine whether the average indoor temperature exceeds the set temperature threshold range; this step is also the above step S212, determining whether the current indoor temperature exceeds the preset temperature range, where the average indoor temperature is also the current indoor temperature, and the set temperature threshold range is also the preset temperature range;
[0086] When the judgment is yes, the action is adjusted according to the traditional control strategy; this step is also the above step S212, if yes, the operation state of the multi-split air conditioner is controlled by the second preset strategy, where the traditional control strategy is also the second preset strategy, that is, the PID control strategy described above;
[0087] When the judgment is no, the step of judging whether the reward value has been improved is entered; this step is also the above step S23, calculating the energy efficiency ratio of the multi-split air conditioner after executing the selected target control action, and comparing it with the saved energy efficiency ratio of the multi-split air conditioner, where the reward value is also the energy efficiency ratio of the multi-split air conditioner;
[0088] When it is determined to be yes again, the process proceeds to update the Q table; this step is also the above-mentioned step S23, in which the larger value in the comparison result is saved, and the process returns to step S21.
[0089] As mentioned above, when the present invention describes the control of the multi-split air conditioner as a reinforcement learning framework, it is necessary to construct the expression model of the current operating state (i.e., the state interval) and the expression model of the control action (i.e., the action interval), and to construct the corresponding calculation model of the energy efficiency ratio of the multi-split air conditioner (i.e., the reward function);
[0090] Specifically, the expression model of the above current running status is:
[0091]
[0092] in, is the observed state of indoor unit i at time t, N is the number of indoor units, is the current indoor temperature of indoor unit i, is the indoor temperature of indoor unit i at the last moment, is the current set temperature of indoor unit i, θ i is the current indoor humidity of indoor unit i, is the current indoor unit capacity of indoor unit i, FL i is the current windshield of indoor unit i, T out is the current outdoor temperature, P o It is the current operating power of the outdoor unit.
[0093] The expression pattern of the above control action is:
[0094] a t ={EXV 1 ,EXV 2 ,L,EXV N ,f com ,f fan}
[0095] Among them, EXV i is the expansion valve opening of unit i at time t, f com is the compressor frequency of the outdoor unit, f fan is the fan frequency.
[0096] The calculation model of the energy efficiency ratio of the above multi-split air conditioner is:
[0097]
[0098] in, is the current internal unit capacity of internal unit i, P o is the current operating power of the outdoor unit, and N is the number of indoor units.
[0099] Among them, in order to balance exploration and utilization, the ε-greedy strategy is used when selecting the control action, which selects the control action through the following formula:
[0100]
[0101] That is, in the current running state t Next, select a control action to make the reward value R t maximum.
[0102] The control method of the multi-split air conditioner proposed by the present invention is described in detail below:
[0103] The reinforcement learning control method is recorded as the RL control strategy, which will only run when the user makes a selection on the user interaction terminal device such as the online controller / remote controller / mobile phone APP. Before entering the RL control strategy, the multi-split air conditioner adjusts the indoor temperature through the traditional PID control strategy;
[0104] It includes step 1: before entering the RL control strategy, such as Figure 3 As shown, first establish the relationship Q table: divide the state space S into n running states, divide the action space A into m control actions, and when executing the control action, record the executed control action a nm The corresponding reward value R nm .
[0105] Step 2: After entering the RL control strategy, monitor the current operating status of the multi-split air conditioner in real time t The current operating state at least includes the current indoor temperature T of each indoor unit. in , the indoor temperature at the last moment T last 、Current user set temperature T set 、Current indoor humidity θ, current indoor unit capacity Q c , current indoor air level FL, current outdoor temperature T out 、Current external machine operating power P o In a preferred embodiment of the present invention, the current operating state is s2, and there are 2 indoor units and 1 outdoor unit in the multi-split air conditioner. At this time, in the relationship Q table, there are 4 control actions corresponding to the current operating state s2, such as Figure 4 As shown;
[0106] Step 3: Based on the ε-greedy strategy, in the relational Q table, the control action corresponding to the current running state s2 selects the target control action and executes it (that is, the attached Figure 4 Then, according to the size of the random number and the set parameters, it enters the greedy strategy or exploration strategy. If it is a greedy strategy, it selects a control action to maximize the reward value R, and then selects the reward value R from the attached Figure 4 It can be seen that the control action selected at this time is a1, and then go to step 4; if it is an exploration strategy, that is, choose to explore the control action that has not been executed, from the attached Figure 4 It can be seen that the control action selected at this time is a2, and then enter step 4;
[0107] Step 3.1: Greedy strategy: Set the initial parameter ε = 0.9. When the generated random number rand∈[0,1] is greater than the initial parameter ε, execute the greedy strategy, otherwise execute the exploration strategy, that is, randomly select a control action that has not been executed;
[0108] Step 3.2: After the control action is selected, ε is updated as follows:
[0109] ε←0.99·ε
[0110] Step 4: After executing the selected control action, calculate the corresponding reward value and determine the current indoor temperature after executing the selected control action. Whether it exceeds the preset temperature range If yes, adjust the control action of the multi-split air conditioner according to the traditional PID control strategy so that the current indoor temperature is within the preset temperature range, and then return to step 1; otherwise, proceed to step 5;
[0111] Step 5: Based on the reward value calculated in the previous step, determine whether the reward value has been improved. For example, when executing the greedy strategy, the latest reward value is 3.61, which is smaller than the original 3.63. In this case, the reward value is not updated and the process returns to step 1. If the latest reward value is 3.65, which is larger than the original 3.63, 3.65 is used instead of 3.63 and the process returns to step 1.
[0112] When executing the exploration strategy, the unexplored control action is the unknown action a2. No matter what the reward value is, the reward value is updated and recorded in Figure 4 and return to step 1.
[0113] The above is the overall control process of the control method of the multi-split air conditioner proposed in the present invention. It can be seen from the above process that the reward value in the present invention is updated in real time, and the final exploration result must satisfy the maximum reward value, that is, the present invention can couple and associate all control variables and enable the multi-split air conditioner to operate at the highest energy efficiency ratio.
[0114] Based on the above control method, the present invention also proposes a multi-split air conditioner, see Figure 1 , which includes an outdoor unit, three indoor units, and an actuator, wherein the actuator is used to execute the control method of the multi-split air conditioner.
[0115] In summary, compared with the prior art, the present invention has at least the following beneficial effects:
[0116] The control method of the multi-split air conditioner proposed in the present invention describes the control of the multi-split air conditioner as a reinforcement learning framework. Both its state space and action space take into account multiple states. At the same time, the energy efficiency ratio of the multi-split air conditioner is used as the final reward value. It can find the optimal control action of the multi-split air conditioner in different user scenarios, which not only takes into account the energy saving of the multi-split air conditioner, but also ensures the comfort of the user.
[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A control method for a multi-split air conditioner, characterized in that: include: Step S1: identifying the current operating state of the multi-split air conditioner, and acquiring all control actions corresponding to the current operating state according to a preset corresponding relationship; Step S2: Based on the first preset strategy, a target control action is selected from all control actions corresponding to the current operating state and executed, and the energy efficiency ratio of the multi-split air conditioner after executing the target control action is calculated in turn, and the target control action corresponding to the highest energy efficiency ratio of the multi-split air conditioner is obtained and executed.
2. The control method of a multi-split air conditioner according to claim 1, characterized in that: The first preset strategy includes a greedy strategy and an exploration strategy; The greedy strategy is used to select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the executed target control actions, and the exploration strategy is used to select the target control action that has not been executed from all the control actions; The greedy strategy and the exploration strategy are switched according to the size relationship between the generated random number and the set parameter.
3. The control method of a multi-split air conditioner according to claim 2, characterized in that: The step S2 comprises: Step S21: Generate a random number and compare it with the set initial parameter. When the random number is smaller than the initial parameter, execute the exploration strategy and enter step S22. When the random number is larger than the initial parameter, execute the greedy strategy and enter step S23. Step S22: Calculate and save the energy efficiency ratio of the multi-split air conditioner after executing the selected target control action, and update the initial parameters at a preset ratio and then return to step S21; Step S23: Select the target control action with the largest energy efficiency ratio of the multi-split air conditioner from the target control actions that have been executed, compare the corresponding energy efficiency ratio of the multi-split air conditioner with the saved energy efficiency ratio of the multi-split air conditioner, save the larger value in the comparison result, and update the initial parameters at a preset ratio before returning to step S21.
4. The control method of a multi-split air conditioner according to claim 3, characterized in that: The step S21 further includes: Step S211: After executing the selected target control action, detecting the current indoor temperature; Step S212: Determine whether the current indoor temperature exceeds the preset temperature range. If so, control the operating state of the multi-split air conditioner with the second preset strategy. If not, execute the greedy strategy or exploration strategy based on the size relationship between the generated random number and the set parameter.
5. The control method of a multi-split air conditioner according to claim 4, characterized in that: The second preset strategy is a PID control strategy.
6. The control method of a multi-split air conditioner according to claim 1, characterized in that: The expression model of the current running state is: in, is the observed state of indoor unit i at time t, N is the number of indoor units, is the current indoor temperature of indoor unit i, is the indoor temperature of indoor unit i at the last moment, is the current set temperature of indoor unit i, θ i is the current indoor humidity of indoor unit i, is the current indoor unit capacity of indoor unit i, FL i is the current windshield of indoor unit i, T out is the current outdoor temperature, P o It is the current operating power of the outdoor unit.
7. The control method of a multi-split air conditioner according to claim 1, characterized in that: The expression model of the control action is: a t ={EXV 1 ,EXV 2 , L, EXV N , f com , f fan } Among them, EXV i is the expansion valve opening of unit i at time t, f com is the compressor frequency of the outdoor unit, f fan is the fan frequency.
8. The control method of a multi-split air conditioner according to claim 1, characterized in that: The calculation model of the energy efficiency ratio of the multi-split air conditioner is: in, is the current internal unit capacity of internal unit i, P o is the current operating power of the outdoor unit, and N is the number of indoor units.
9. The control method of a multi-split air conditioner according to claim 3, characterized in that: The initial parameter is set to 0.9, the random number is selected from the interval between 0 and 1, and the preset ratio is 0.
99.
10. A multi-split air conditioner, characterized in that: The multi-split air conditioner has one outdoor unit, three indoor units, and an actuator, and the actuator is used to execute the control method of the multi-split air conditioner according to any one of claims 1 to 9.