Multi-connected air conditioner control method, device, equipment, storage medium and product
By acquiring the state-space parameters of a multi-split air conditioner and training a basic control model with the goal of minimizing power consumption, the problem of multi-split air conditioners being unable to achieve precise temperature control and energy saving at the same time is solved, thus achieving precise temperature control while reducing energy consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAOMI TECH (WUHAN) CO LTD
- Filing Date
- 2024-12-05
- Publication Date
- 2026-06-05
AI Technical Summary
Multi-split air conditioners have difficulty achieving precise temperature control and energy saving simultaneously during the control process.
By acquiring the state-space parameters of the outdoor and indoor units, processing them using a target control model, training the basic control model with the goal of minimizing power consumption, obtaining the control parameters of the outdoor and indoor units, and then controlling them based on these parameters.
It achieves precise temperature control while reducing energy consumption, and improves control accuracy and energy-saving effect by learning the collaborative relationship between indoor and outdoor units through centralized training.
Smart Images

Figure CN122149052A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of air conditioning control technology, and in particular to a method, apparatus, equipment, storage medium and product for controlling multi-split air conditioning systems. Background Technology
[0002] Controlling multi-split air conditioners requires coordinated control of the outdoor unit and multiple indoor units, making it difficult to achieve precise temperature control and energy saving simultaneously. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides a multi-split air conditioning control method, device, equipment, storage medium and product, so as to reduce energy consumption while achieving precise temperature control.
[0004] According to a first aspect of the present disclosure, a method for controlling a multi-split air conditioner is provided, the multi-split air conditioner including an outdoor unit and multiple indoor units, the method comprising: The first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are obtained. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, indoor unit operating parameters and indoor unit environment parameters of the corresponding indoor unit. The first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are processed by the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is based on the simulation environment corresponding to the multi-split air conditioner, with the minimum power consumption as the training objective, and the basic control model is trained to obtain the target control model. The outdoor unit is controlled based on the first control parameter, and the corresponding indoor unit is controlled based on the second control parameter of each indoor unit.
[0005] Optionally, the target control model includes multiple distributed sub-models; The process of processing the first state-space parameters of the outdoor unit and the second state-space parameters of each indoor unit using the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit includes: Determine the distributed sub-model corresponding to each external machine and each internal machine, with one external machine or one internal machine corresponding to one distributed sub-model; The first state space parameters of the outdoor unit are input into the distributed sub-model corresponding to the outdoor unit to obtain the first control parameters of the outdoor unit; The second state space parameters of each internal unit are input into the distributed sub-model corresponding to each internal unit to obtain the second control parameters of each internal unit.
[0006] Optionally, the target control model is trained through the following steps: Based on the multi-split air conditioner, a corresponding simulation environment is constructed, which includes a simulated outdoor unit and multiple simulated indoor units; The basic control model, including a total reward network and a control network, is trained through multiple rounds of iterative training in the simulation environment. After each round of training, determine the error loss corresponding to that round of training; Based on the aforementioned error loss, the basic control model is optimized; If the basic control model meets the preset conditions, training is stopped, and the target control model is obtained.
[0007] Optionally, the control network includes multiple distributed sub-networks, with one simulated outdoor unit or one simulated indoor unit corresponding to one distributed sub-network; The step of performing multiple rounds of iterative training on the basic control model in the simulation environment includes: For each round of iterative training, the first simulated state space parameters corresponding to this round of training are obtained through the simulated environment. The first simulated state space parameters include the state space parameters corresponding to the simulated outdoor unit and the state space parameters corresponding to each simulated indoor unit. The state space parameters corresponding to the simulated outdoor unit are input into the distributed sub-network corresponding to the simulated outdoor unit to obtain the first simulated reward parameter of the simulated outdoor unit, wherein the first simulated reward parameter includes the first simulated control parameter. The state space parameters corresponding to each simulated intranet are input into the distributed subnetwork corresponding to each simulated intranet to obtain the second simulated reward parameters of each simulated intranet. The second simulated reward parameters include the second simulated control parameters.
[0008] Optionally, determining the error loss corresponding to each training round after each round includes: After each round of training, obtain the first simulated reward parameters and the second simulated reward parameters corresponding to this round of training; Input the first and second simulation control parameters corresponding to this round of training into the simulation environment to obtain the power consumption parameters corresponding to this round of training and the second simulation state space parameters corresponding to the next round of training. Based on the first simulated control parameters, the second simulated control parameters, the power consumption parameters, and the control constraints, the reward value corresponding to this round of training is obtained; Input the first simulated reward parameter, the second simulated reward parameter, and the second simulated state space parameter into the total reward network to obtain the estimated total reward for this round of training. The error loss for this training round is obtained based on the first simulated state space parameters, the second simulated state space parameters, the reward value and estimated total reward for this training round, and the estimated total reward for the previous training round.
[0009] Optionally, obtaining the reward value corresponding to this round of training based on the first simulation control parameters, the second simulation control parameters, the power consumption parameters, and the control constraints includes: If both the first and second simulated control parameters satisfy the control constraints, the reward value corresponding to this round of training is obtained based on the power consumption parameter. If either the first simulation control parameter or the second simulation control parameter fails to meet the control constraint, the reward value corresponding to this round of training is obtained based on the maximum reward threshold.
[0010] Optionally, the power consumption parameters include actual power consumption, execution time, set temperature, and indoor ambient temperature; The step of obtaining the reward value corresponding to this round of training based on the power consumption parameter includes: The reward value corresponding to the simulated outdoor unit is determined based on the actual power consumption of the simulated outdoor unit. The reward value for each simulated indoor unit is determined based on its actual power consumption, execution time, set temperature, and indoor ambient temperature. The sum of the reward value corresponding to the simulated outdoor unit and the reward value corresponding to each simulated indoor unit is determined as the reward value for this round of training.
[0011] Optionally, determining the reward value for each simulated indoor unit based on its actual power consumption, execution time, set temperature, and indoor ambient temperature includes: For any simulated indoor unit in each simulated indoor unit, if the indoor ambient temperature corresponding to the simulated indoor unit is not equal to the set temperature, the correction parameters corresponding to the simulated indoor unit are obtained based on the execution time, set temperature and indoor ambient temperature corresponding to the simulated indoor unit. The reward value corresponding to the simulated indoor unit is obtained based on the correction parameters corresponding to the simulated indoor unit and the actual power consumption. When the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature, the reward value corresponding to the simulated indoor unit is obtained by subtracting the actual power consumption from the maximum reward threshold.
[0012] Optionally, the total reward network includes a parameter generation network and a combination network; The step of inputting the first simulated reward parameter, the second simulated reward parameter, and the second simulated state space parameter into the total reward network to obtain the estimated total reward corresponding to this round of training includes: The second simulated state-space parameters are processed by the parameter generation network to obtain a first combined parameter, a second combined parameter, a first bias term, and a second bias term; Through the combined network, the first combined parameter, the first bias term, the second combined parameter, and the second bias term are sequentially combined with the first simulated reward parameter and the second simulated reward parameter to obtain the estimated total reward corresponding to this round of training.
[0013] Optionally, the method further includes: Obtain the historical state space parameters corresponding to the multi-split air conditioner within the target time period; The target control model is updated based on the historical state space parameters to obtain the updated target control model.
[0014] According to a second aspect of the present disclosure, a multi-split air conditioner control device is provided, configured in a multi-split air conditioner, the multi-split air conditioner including an outdoor unit and a plurality of indoor units, the device comprising: The first acquisition module is configured to acquire the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, the indoor unit operating parameters and the indoor unit environment parameters of the corresponding indoor unit. The first acquisition module is configured to process the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit through the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is obtained by training the basic control model based on the simulation environment corresponding to the multi-split air conditioner with the minimum power consumption as the training objective. The control module is configured to control the outdoor unit based on the first control parameter, and to control the corresponding indoor unit based on the second control parameter of each indoor unit.
[0015] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to execute the steps of the multi-split air conditioning control method provided in the first aspect of this disclosure.
[0016] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the multi-split air conditioning control method provided in the first aspect of the present disclosure.
[0017] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the multi-split air conditioning control method provided in the first aspect of the present disclosure.
[0018] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: The system acquires the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit. The first state space parameters include the outdoor unit's operating parameters and environmental parameters, while the second state space parameters include the corresponding indoor unit's setting parameters, operating parameters, and environmental parameters. A target control model is then used to process these parameters to obtain the first control parameters for the outdoor unit and the second control parameters for each indoor unit. The target control model is trained based on a simulated environment of the multi-split air conditioner, with the goal of minimizing power consumption. Finally, the system controls the outdoor unit based on the first control parameters and the corresponding indoor unit based on its second control parameters.
[0019] Based on the simulation environment corresponding to multi-split air conditioners, the entire system of a multi-split air conditioner can be simulated. The basic control model is trained with the goal of minimizing power consumption, achieving a centralized training effect. This allows the basic control model to learn the collaborative relationships between multiple indoor units, as well as the collaborative relationships between multiple indoor units and the outdoor unit, resulting in a target control model. Furthermore, by processing the first state-space parameters of the outdoor unit and the second state-space parameters of each indoor unit using the target control model, more accurate first control parameters for the outdoor unit and second control parameters for each indoor unit can be obtained. This enables precise temperature control while reducing energy consumption.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of a multi-split air conditioning control method according to an exemplary embodiment.
[0023] Figure 2 This is a flowchart illustrating a multi-split air conditioning control method according to an exemplary embodiment.
[0024] Figure 3 This is a schematic diagram illustrating a model training according to an exemplary embodiment.
[0025] Figure 4 This is a block diagram illustrating a multi-split air conditioning control device according to an exemplary embodiment.
[0026] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements.
[0028] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0029] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0030] Controlling multi-split air conditioners requires coordinated control of the outdoor unit and multiple indoor units, making it difficult to achieve precise temperature control and energy saving simultaneously.
[0031] To address the aforementioned technical problems, this disclosure provides a multi-split air conditioner control method, apparatus, device, storage medium, and product. Based on a simulation environment corresponding to the multi-split air conditioner, it can simulate the entire system of the multi-split air conditioner. With minimizing power consumption as the training objective, a basic control model is trained, achieving a centralized training effect. This allows the basic control model to learn the collaborative relationships between multiple indoor units and between multiple indoor units and the outdoor unit, resulting in a target control model. Furthermore, by processing the first state-space parameters of the outdoor unit and the second state-space parameters of each indoor unit using the target control model, more accurate first control parameters for the outdoor unit and second control parameters for each indoor unit can be obtained, thereby achieving precise temperature control while reducing energy consumption.
[0032] Figure 1 This is a schematic diagram illustrating an application scenario of a multi-split air conditioning control method according to an exemplary embodiment, such as... Figure 1As shown, this can be applied to a multi-split air conditioning system, which may include one outdoor unit and multiple indoor units. The indoor units are installed in different rooms within the building, while the outdoor units are located outdoors. Each indoor unit has corresponding indoor unit setting parameters, indoor unit operating parameters, and indoor unit environmental parameters, and each outdoor unit has corresponding outdoor unit operating parameters and outdoor unit environmental parameters. Additionally, the multi-split air conditioning system may include a controller, which can be located on one of the indoor or outdoor units, or it can be a separate controller that can communicate with each indoor and outdoor unit. Furthermore, the multi-split air conditioning system may also include a cloud data center, which can communicate with each indoor and outdoor unit.
[0033] Figure 2 This is a flowchart illustrating a multi-split air conditioning control method according to an exemplary embodiment, such as... Figure 2 As shown, this method can be applied to multi-split air conditioners, controllers, or cloud data centers. The multi-split air conditioner may include one outdoor unit and multiple indoor units. The method may include the following steps.
[0034] In step S201, the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are obtained. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, indoor unit operating parameters and indoor unit environment parameters of the corresponding indoor unit.
[0035] In this embodiment, the first state space parameters corresponding to the outdoor unit can be obtained directly from the outdoor unit, and the second state space parameters corresponding to each indoor unit can be obtained directly from each indoor unit; alternatively, the first state space parameters corresponding to the outdoor unit can be uploaded through the outdoor unit, and the second state space parameters corresponding to each indoor unit can be uploaded through each indoor unit. The first state space parameters may include outdoor unit operating parameters and outdoor unit environmental parameters. The outdoor unit operating parameters may include at least one of compressor frequency and outdoor fan speed. The outdoor unit environmental parameters may include at least one of the duct length between each indoor unit and the outdoor unit, outdoor ambient temperature, compressor exhaust temperature, compressor inlet pressure, compressor outlet pressure, height difference between each indoor unit and the outdoor unit, and room height of the room where each indoor unit is located. The second state space parameters may include the indoor unit setting parameters, indoor unit operating parameters, and indoor unit environmental parameters. The indoor unit setting parameters may be the set temperature of the indoor unit; the indoor unit operating parameters may be the indoor fan speed and expansion valve opening; and the indoor unit environmental parameters may be at least one of the indoor ambient humidity, room area, indoor ambient temperature, height difference between the indoor and outdoor units, and room height of the area where the indoor unit is located. By obtaining these parameters, we can obtain more accurate air conditioning control parameters.
[0036] In step S202, the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are processed by the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is based on the simulation environment corresponding to the multi-split air conditioner, with the minimum power consumption as the training objective, and the basic control model is trained.
[0037] In this embodiment, a basic control model can be pre-trained in a simulated environment corresponding to the multi-split air conditioner to obtain a target control model. During training, the goal is to minimize power consumption. If the basic control model is a reinforcement learning model, the reward can be determined based on power consumption. The target control model is then processed to obtain the first control parameters for the outdoor unit and the second control parameters for each indoor unit. The first control parameters may include the compressor frequency and outdoor fan speed of the outdoor unit, while the second control parameters may include the expansion valve opening and indoor fan speed of the indoor unit.
[0038] Based on the simulation environment corresponding to multi-split air conditioners, the entire system of a multi-split air conditioner can be simulated. The basic control model is trained with the goal of minimizing power consumption, achieving a centralized training effect. This allows the basic control model to learn the collaborative relationships between multiple indoor units, as well as the collaborative relationships between multiple indoor units and the outdoor unit, resulting in a target control model. Furthermore, by processing the first state-space parameters of the outdoor unit and the second state-space parameters of each indoor unit using the target control model, more accurate first control parameters for the outdoor unit and second control parameters for each indoor unit can be obtained. This enables precise temperature control while reducing energy consumption.
[0039] In step S203, the external unit is controlled based on the first control parameter, and the corresponding internal unit is controlled based on the second control parameter of each internal unit.
[0040] In this embodiment, after obtaining the first control parameters of the outdoor unit and the second control parameters of each indoor unit, the indoor unit can be controlled by the first control parameters, and the corresponding indoor unit can be controlled by the second control parameters corresponding to each indoor unit, thereby reducing energy consumption while achieving precise temperature control.
[0041] In one possible implementation, the target control model may include multiple distributed sub-models, with one distributed sub-model corresponding to one external unit or one internal unit. These distributed sub-models may be set up separately in their respective internal or external units, or they may be uniformly set up within the controller.
[0042] By processing the first state-space parameters of the external unit and the second state-space parameters of each internal unit using the target control model, the first control parameters of the external unit and the second control parameters of each internal unit are obtained, including: Determine the distributed sub-models corresponding to the outdoor unit and each indoor unit respectively, with one distributed sub-model corresponding to one outdoor unit or one indoor unit; input the first state space parameters of the outdoor unit into the distributed sub-model corresponding to the outdoor unit to obtain the first control parameters of the outdoor unit; input the second state space parameters of each indoor unit into the distributed sub-model corresponding to each indoor unit to obtain the second control parameters of each indoor unit.
[0043] In this embodiment, the distributed sub-models corresponding to the outdoor unit and each indoor unit can be determined first. If multiple distributed sub-models are uniformly set within the controller, the distributed sub-models corresponding to the outdoor unit and each indoor unit can be determined based on the correspondence between each indoor unit or outdoor unit and the distributed sub-model. If multiple distributed sub-models can be set within their respective indoor or outdoor units, the distributed sub-model set within each indoor or outdoor unit is determined as its corresponding distributed sub-model. Furthermore, for any indoor unit, the second state space parameter corresponding to that indoor unit can be input into the distributed sub-model corresponding to that indoor unit to obtain the second control parameter corresponding to that indoor unit. For an outdoor unit, the first state space parameter corresponding to that outdoor unit can be input into the distributed sub-model corresponding to that outdoor unit to obtain the first control parameter corresponding to that outdoor unit. Using the above method, the corresponding parameters can be processed through the distributed sub-models of the indoor and outdoor units respectively, thereby obtaining the corresponding control parameters. This reduces the overall data processing pressure. Moreover, each indoor and outdoor unit only needs to obtain its own corresponding control parameters based on its own parameters, which meets the overall requirements for precise temperature control and energy saving. It does not need to obtain the parameters of other indoor or outdoor units, thereby reducing data transmission, lowering the pressure of data transmission, and avoiding data errors that may occur during data transmission.
[0044] In one possible implementation, the target control model is trained through the following steps: Based on a multi-split air conditioner, a corresponding simulation environment is constructed, which includes one simulated outdoor unit and multiple simulated indoor units. Through the simulation environment, the basic control model is trained iteratively in multiple rounds. The basic control model includes a total reward network and a control network. After each round of training, the error loss corresponding to this round of training is determined. Based on the error loss, the basic control model is optimized. When the basic control model meets the preset conditions, training is stopped, and the target control model is obtained.
[0045] In this embodiment, a corresponding simulation environment can be constructed based on a multi-split air conditioner. Specifically, based on the number of indoor and outdoor units in the multi-split air conditioner, the performance parameters of each indoor and outdoor unit, and combined with sample setting parameters and sample environment parameters, a simulation environment that can simulate the operation of the multi-split air conditioner under the sample environment can be constructed. Furthermore, the basic control model can be trained iteratively through the simulation environment. The basic control model may include a total reward network and a control network. The control network is used to output simulated control parameters and simulated reward parameters, and the total reward network is used to output the estimated total reward.
[0046] After each training round, the error loss corresponding to this round of training can be calculated based on the output parameters. Based on this error loss, optimized parameters for the basic control model can be obtained through backpropagation. These optimized parameters may include those for the total reward network and the control network, aiming to make their outputs more accurate. Training can stop when the basic control model meets preset conditions, resulting in the target control model. These preset conditions may include reaching the target number of training rounds, achieving the performance target, or ceasing further improvement.
[0047] The control network can include multiple distributed sub-networks. Each distributed sub-network can include two MLP (Multilayer Perceptron) layers and one GRU (Gated Recurrent Unit) layer. After the target control model is trained, the control network in the target control model can be allocated to the actual multi-split air conditioner. One indoor or outdoor unit of a multi-split air conditioner corresponds to one distributed sub-network, thus obtaining the distributed sub-model corresponding to the indoor or outdoor unit.
[0048] Figure 3 This is a schematic diagram illustrating a model training method according to an exemplary embodiment, such as... Figure 3 As shown, in one possible implementation, the control network includes multiple distributed sub-networks, with one simulated outdoor unit or one simulated indoor unit corresponding to one distributed sub-network; the basic control model is trained iteratively through a simulated environment, including: For each round of iterative training, the first simulated state space parameters corresponding to this round of training are obtained through the simulated environment. The first simulated state space parameters include the state space parameters corresponding to the simulated external machine and the state space parameters corresponding to each simulated internal machine. The state space parameters corresponding to the simulated external machine are input into the distributed sub-network corresponding to the simulated external machine to obtain the first simulated reward parameters of the simulated external machine. The first simulated reward parameters include the first simulated control parameters. The state space parameters corresponding to each simulated internal machine are input into the distributed sub-network corresponding to each simulated internal machine to obtain the second simulated reward parameters of each simulated internal machine. The second simulated reward parameters include the second simulated control parameters.
[0049] In this embodiment, a simulation environment can be run to obtain the first simulation state space parameters corresponding to this round of training. These first simulation state space parameters may include the state space parameters corresponding to the simulated outdoor unit and the state space parameters corresponding to each simulated indoor unit. Specifically, the state space parameters corresponding to the simulated outdoor unit may include outdoor unit operating parameters and outdoor unit environmental parameters, while the state space parameters corresponding to the simulated indoor unit may include the indoor unit setting parameters, indoor unit operating parameters, and indoor unit environmental parameters. The outdoor unit operating parameters may include at least one of the compressor frequency and the outdoor fan speed, and the outdoor unit environmental parameters may include at least one of the following: the duct length between each simulated indoor unit and the simulated outdoor unit, the outdoor ambient temperature, the compressor exhaust temperature, the compressor inlet pressure, the compressor outlet pressure, the height difference between each simulated indoor unit and the simulated outdoor unit, and the room height of the room where each simulated indoor unit is located. Among them, the indoor unit setting parameters, i.e. the sample setting parameters, can be the sample setting temperature of the simulated indoor unit; the indoor unit operating parameters can be the indoor fan speed and expansion valve opening of the simulated indoor unit; the indoor unit environmental parameters can be at least one of the indoor environmental humidity, room area, indoor environmental temperature, indoor-outdoor unit height difference, and room floor height corresponding to the area where the simulated indoor unit is located.
[0050] By inputting the state space parameters corresponding to the simulated outdoor unit into the distributed sub-network corresponding to the simulated outdoor unit, the first simulated reward parameters of the simulated outdoor unit can be obtained. The first simulated reward parameters include the first simulated control parameters. The state space parameters corresponding to each simulated indoor unit are then input into the distributed sub-network corresponding to each simulated indoor unit to obtain the second simulated reward parameters of each simulated indoor unit. The second simulated reward parameters include the second simulated control parameters. The first simulated control parameters may include the compressor frequency and outdoor fan speed corresponding to the simulated outdoor unit, while the second simulated control parameters may include the expansion valve opening and indoor fan speed corresponding to the simulated indoor unit. This allows for the control of the simulated indoor and outdoor units in the simulated environment based on the first and second simulated control parameters, resulting in further second simulated state space parameters. This achieves the effect of centralized training and distributed execution, enabling the learning of the collaborative relationships between multiple indoor units and between multiple indoor and outdoor units.
[0051] In one possible implementation, after each round of training, the error loss corresponding to that round of training is determined, including: After each training round, the first and second simulated reward parameters corresponding to this training round are obtained; the first and second simulated control parameters corresponding to this training round are input into the simulation environment to obtain the power consumption parameters corresponding to this training round and the second simulated state space parameters corresponding to the next training round; based on the first and second simulated control parameters, the power consumption parameters, and the control constraints, the reward value corresponding to this training round is obtained; the first, second, and second simulated reward parameters are input into the total reward network to obtain the estimated total reward corresponding to this training round; based on the first and second simulated state space parameters, the reward value corresponding to this training round, the estimated total reward, and the estimated total reward corresponding to the previous training round, the error loss corresponding to this training round is obtained.
[0052] In this embodiment, each training round may include a target number of steps. After reaching the target number of steps, the error loss can be calculated to update the basic control model. Specifically, after each training round, all the first and second simulated reward parameters corresponding to this round of training can be obtained. The first and second simulated control parameters from the first and second simulated reward parameters can be input into the simulation environment and run for a certain period of time to obtain the power consumption parameters corresponding to this round of training and the second simulated state space parameters corresponding to the next round of training. The specific parameters included in the second simulated state space parameters corresponding to the next round of training can be the same as the parameters included in the first simulated state space.
[0053] Then, based on the reward value calculation method, the reward value corresponding to this round of training can be obtained according to the first simulated control parameters, the second simulated control parameters, the power consumption parameters, and the control constraints. The first simulated reward parameters, the second simulated reward parameters, and the second simulated state space parameters are then input into the total reward network to obtain the estimated total reward for this round of training. Combining this with the error loss calculation formula, the error loss for this round of training can be obtained based on the first simulated state space parameters, the second simulated state space parameters, the reward value and estimated total reward for this round of training, and the estimated total reward for the previous round of training. The control constraints may include compressor frequency range, external fan speed range, internal fan speed range, and expansion valve opening range, to constrain the compressor frequency, external fan speed, internal fan speed, and expansion valve opening, thus preventing the model from slowing down its learning process to conserve power during training.
[0054] The error loss can be the timing difference error, and the corresponding calculation formula is as follows:
[0055] in, For error loss, This is the reward value corresponding to this round of training. For the first Rotation training, It is an adjustable constant value. This is the estimated total reward for this round of training. This is the estimated total reward corresponding to the previous round of training. These are the parameters of the second simulated state space. These are the parameters of the first simulated state space.
[0056] In one possible implementation, the reward value corresponding to this round of training is obtained based on the first simulated control parameters, the second simulated control parameters, the power consumption parameters, and the control constraints, including: If both the first and second simulated control parameters satisfy the control constraints, the reward value for this round of training is obtained based on the power consumption parameter; if either the first or second simulated control parameter fails to satisfy the control constraints, the reward value for this round of training is obtained based on the maximum reward threshold.
[0057] In this embodiment, if both the first and second simulated control parameters satisfy the control constraints, no penalty is required, and the reward value can be calculated normally based on power consumption, i.e., the reward value for this round of training is obtained based on the power consumption parameter. If either the first or second simulated control parameter fails to satisfy the control constraints, the learning direction is considered incorrect, and the reward value for this round of training can be obtained based on the maximum reward threshold. Specifically, the negative of the maximum reward threshold can be determined as the reward value for this round of training, i.e., a larger penalty value is directly applied to adjust the direction of model learning.
[0058] The formula for calculating the reward value for this round of training can be:
[0059] in, The maximum reward threshold, This refers to the serial number corresponding to either the simulated indoor unit or the simulated outdoor unit. When the value is 0, it corresponds to the simulated outdoor unit. When the value is not equal to 0, it corresponds to the simulated indoor unit.
[0060] In one possible implementation, the power consumption parameters include actual power consumption, execution time, set temperature, and indoor ambient temperature.
[0061] The reward value for this round of training is obtained based on the power consumption parameters, including: The reward value for each simulated outdoor unit is determined based on its actual power consumption. The reward value for each simulated indoor unit is determined based on its actual power consumption, execution time, set temperature, and indoor ambient temperature. The sum of the reward values for the simulated outdoor units and the reward values for each simulated indoor unit is determined as the reward value for this round of training.
[0062] In this implementation, for the simulated outdoor unit, its reward value is determined directly based on its power consumption. For each simulated indoor unit, its reward value is determined by considering its actual power consumption and whether it has reached the set temperature. After obtaining the reward values for each simulated indoor and outdoor unit, they are summed to obtain the total reward value, which is then used as the reward value for this round of training.
[0063] In one possible implementation, the reward value for each simulated indoor unit is determined based on its actual power consumption, execution time, set temperature, and indoor ambient temperature, including: For any simulated indoor unit in each simulated indoor unit, if the indoor ambient temperature corresponding to the simulated indoor unit is not equal to the set temperature, the correction parameters corresponding to the simulated indoor unit are obtained based on the execution time, set temperature, and indoor ambient temperature of the simulated indoor unit; the reward value corresponding to the simulated indoor unit is obtained based on the correction parameters corresponding to the simulated indoor unit and the actual power consumption; if the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature, the maximum reward threshold is subtracted from the actual power consumption to obtain the reward value corresponding to the simulated indoor unit.
[0064] In this embodiment, for any simulated indoor unit, it is first determined whether the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature, which is the sample set temperature mentioned earlier. If the indoor ambient temperature corresponding to the simulated indoor unit is not equal to the set temperature, the learning direction needs to be further adjusted. Based on the execution time, set temperature, and indoor ambient temperature of the simulated indoor unit, correction parameters are obtained to achieve precise temperature control and reduce power consumption. If the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature, the learning direction is correct, and a maximum reward threshold can be assigned. Subtracting the actual power consumption from the maximum reward threshold yields the reward value for the simulated indoor unit. This ensures that the learned model can achieve precise temperature control and reduce power consumption.
[0065] The formula for calculating the reward value for each simulated machine can be as follows:
[0066] in, This represents the reward value corresponding to the i-th simulated internal unit. Let i be the actual power consumption of the i-th simulated indoor unit. Let i be the indoor ambient temperature corresponding to the i-th simulated indoor unit. This is the set temperature corresponding to the i-th simulated indoor unit.
[0067] In one possible implementation, the total reward network includes a parameter generation network and a combination network.
[0068] Inputting the first simulated reward parameters, the second simulated reward parameters, and the second simulated state space parameters into the total reward network yields the estimated total reward for this training round, including: The parameters of the second simulated state space are processed by a parameter generation network to obtain a first combined parameter, a second combined parameter, a first bias term, and a second bias term. The first combined parameter, the first bias term, the second combined parameter, and the second bias term are then combined with the first simulated reward parameter and the second simulated reward parameter in sequence by a combination network to obtain the estimated total reward corresponding to this round of training.
[0069] In this embodiment, the total reward network can perform unified data processing, enabling the model to learn an estimated total reward value. This allows the model to learn the collaboration between the simulated internal and external machines, as well as the collaboration between individual simulated internal machines. The total reward network may include a parameter generation network and a combination network. Based on the second simulated state space parameters, the total reward network can obtain corresponding first combination parameters, second combination parameters, first bias terms, and second bias terms. Then, it combines these parameters with the first and second simulated reward parameters corresponding to the simulated internal and external machines through the combination network to obtain the estimated total reward for this round of training. The parameter generation network may include a first generation network for generating the first combination parameters, a second generation network for generating the second combination parameters, a third generation network for generating the first bias term, and a fourth generation network for generating the second bias term. The first, second, and fourth generation networks each include two linear layers and one ReLU (Rectified Linear Unit) layer for generation, while the third generation network includes one linear layer.
[0070] In one possible implementation, the method further includes: Obtain the historical state space parameters of the multi-split air conditioner within the target time period; update the target control model based on the historical state space parameters to obtain the updated target control model.
[0071] In this embodiment, historical state-space parameters corresponding to the multi-split air conditioner within the target time period can be obtained and used as training samples to train the target control model. This allows for updating the target control model and enables online learning, ensuring that the output control parameters better match the actual situation, achieving more precise temperature control and reduced energy consumption. The training process can refer to the training process for obtaining the target control model described above, and will not be repeated here. The target time period can be from 3 AM to 12 AM every day.
[0072] Figure 4 This is a block diagram illustrating a multi-split air conditioning control device according to an exemplary embodiment. (Refer to...) Figure 4 The multi-split air conditioner control device 400 is configured in a multi-split air conditioner, which includes an outdoor unit and multiple indoor units. The multi-split air conditioner control device 400 includes a first acquisition module 401, a first acquisition module 402, and a control module 403.
[0073] The first acquisition module 401 is configured to acquire the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, the indoor unit operating parameters and the indoor unit environment parameters of the corresponding indoor unit. The first acquisition module 402 is configured to process the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit through the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is based on the simulation environment corresponding to the multi-split air conditioner, with the minimum power consumption as the training objective, and is obtained by training the basic control model. The control module 403 is configured to control the outdoor unit based on the first control parameters, and to control the corresponding indoor unit based on the second control parameters of each indoor unit.
[0074] Optionally, the target control model includes multiple distributed sub-models; The first obtaining module 402 includes: The determination submodule is configured to determine the distributed sub-model corresponding to each external machine and each internal machine, with one external machine or one internal machine corresponding to one distributed sub-model. The first acquisition submodule is configured to input the first state space parameters of the external machine into the distributed sub-model corresponding to the external machine to obtain the first control parameters of the external machine. The second acquisition submodule is configured to input the second state space parameters of each internal unit into the distributed sub-model corresponding to each internal unit to obtain the second control parameters of each internal unit.
[0075] Optionally, the multi-split air conditioning control device 400 further includes: The construction module is configured to build a corresponding simulation environment based on the multi-split air conditioner, the simulation environment including a simulated outdoor unit and multiple simulated indoor units; The training module is configured to perform multiple rounds of iterative training on the basic control model through the simulation environment, the basic control model including a total reward network and a control network; The determination module is configured to determine the error loss corresponding to each training round after each training round. The optimization module is configured to optimize the basic control model based on the error loss; The second acquisition module is configured to stop training and obtain the target control model when the basic control model meets preset conditions.
[0076] Optionally, the control network includes multiple distributed sub-networks, with one simulated outdoor unit or one simulated indoor unit corresponding to one distributed sub-network; The training module includes: The first acquisition submodule is configured to acquire the first simulated state space parameters corresponding to the current training round through the simulated environment for each round of iterative training. The first simulated state space parameters include the state space parameters corresponding to the simulated outdoor unit and the state space parameters corresponding to each simulated indoor unit. The third acquisition submodule is configured to input the state space parameters corresponding to the simulated outdoor unit into the distributed subnetwork corresponding to the simulated outdoor unit to obtain the first simulation reward parameter of the simulated outdoor unit, wherein the first simulation reward parameter includes the first simulation control parameter. The fourth submodule is configured to input the state space parameters corresponding to each simulated intranet into the distributed subnetwork corresponding to each simulated intranet to obtain the second simulated reward parameters of each simulated intranet, wherein the second simulated reward parameters include the second simulated control parameters.
[0077] Optionally, the determining module includes: The second acquisition submodule is configured to acquire the first simulated reward parameter and the second simulated reward parameter corresponding to the current training round after each training round. The fifth submodule is configured to input the first and second simulation control parameters corresponding to the current training round into the simulation environment to obtain the power consumption parameters corresponding to the current training round and the second simulation state space parameters corresponding to the next training round. The sixth submodule is configured to obtain the reward value corresponding to this round of training based on the first simulation control parameters, the second simulation control parameters, the power consumption parameters, and the control constraints. The seventh submodule is configured to input the first simulated reward parameter, the second simulated reward parameter, and the second simulated state space parameter into the total reward network to obtain the estimated total reward corresponding to this round of training. The eighth submodule is configured to obtain the error loss corresponding to the current training round based on the first simulated state space parameters, the second simulated state space parameters, the reward value and estimated total reward corresponding to the current training round, and the estimated total reward corresponding to the previous training round.
[0078] Optionally, the sixth obtaining submodule includes: The first obtaining unit is configured to obtain the reward value corresponding to this round of training based on the power consumption parameter, provided that both the first simulation control parameter and the second simulation control parameter satisfy the control constraint condition. The second obtaining unit is configured to obtain the reward value corresponding to the current training round based on the maximum reward threshold if either the first simulation control parameter or the second simulation control parameter fails to meet the control constraint condition.
[0079] Optionally, the power consumption parameters include actual power consumption, execution time, set temperature, and indoor ambient temperature; The first obtaining unit includes: The first determining subunit is configured to determine the reward value corresponding to the simulated outdoor unit based on the actual power consumption of the simulated outdoor unit. The second determining subunit is configured to determine the reward value for each simulated indoor unit based on the actual power consumption, execution time, set temperature and indoor ambient temperature of each simulated indoor unit. The third determining subunit is configured to determine the sum of the reward value corresponding to the simulated outdoor unit and the reward value corresponding to each simulated indoor unit as the reward value corresponding to this round of training.
[0080] Optionally, the second determining subunit includes: The first obtaining node is configured to, for any simulated indoor unit in each simulated indoor unit, if the indoor ambient temperature corresponding to the simulated indoor unit is not equal to the set temperature, obtain the correction parameters corresponding to the simulated indoor unit based on the execution time, set temperature and indoor ambient temperature corresponding to the simulated indoor unit; The second obtaining node is configured to obtain the reward value corresponding to the simulated indoor unit based on the correction parameters corresponding to the simulated indoor unit and the actual power consumption; The third obtaining node is configured to subtract the actual power consumption from the maximum reward threshold to obtain the reward value corresponding to the simulated indoor unit when the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature.
[0081] Optionally, the total reward network includes a parameter generation network and a combination network; The seventh obtaining submodule includes: The third obtaining unit is configured to process the second simulated state space parameters through the parameter generation network to obtain a first combined parameter, a second combined parameter, a first bias term, and a second bias term; The fourth obtaining unit is configured to combine the first combining parameter, the first bias term, the second combining parameter, and the second bias term with the first simulated reward parameter and the second simulated reward parameter in sequence through the combining network to obtain the estimated total reward corresponding to this round of training.
[0082] Optionally, the multi-split air conditioning control device 400 further includes: The second acquisition module is configured to acquire the historical state space parameters corresponding to the multi-split air conditioner within the target time period. The update module is configured to update the target control model based on the historical state space parameters to obtain the updated target control model.
[0083] Regarding the multi-split air conditioning control device 400 in the above embodiments, the specific methods by which each module performs its operation have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0084] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the multi-split air conditioning control method provided in this disclosure.
[0085] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment. For example, the electronic device 500 may be a multi-split air conditioner or an air conditioner controller.
[0086] Reference Figure 5 The electronic device 500 may include one or more of the following components: a first processing component 502, a first memory 504, a first power supply component 506, a multimedia component 508, an audio component 510, a first input / output interface 512, a sensor component 514, and a communication component 516.
[0087] The first processing component 502 typically controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. The first processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the multi-split air conditioning control method described above. Furthermore, the first processing component 502 may include one or more modules to facilitate interaction between the first processing component 502 and other components. For example, the first processing component 502 may include a multimedia module to facilitate interaction between the multimedia component 508 and the first processing component 502.
[0088] The first memory 504 is configured to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phonebook data, messages, pictures, videos, etc. The first memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0089] The first power supply component 506 provides power to various components of the electronic device 500. The first power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 500.
[0090] Multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0091] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when electronic device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in first memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.
[0092] The first input / output interface 512 provides an interface between the first processing component 502 and the peripheral interface module, which may be a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.
[0093] Sensor assembly 514 includes one or more sensors for providing state assessments of various aspects of electronic device 500. For example, sensor assembly 514 may detect the on / off state of electronic device 500, the relative positioning of components such as the display and keypad of electronic device 500, changes in position of electronic device 500 or a component of electronic device 500, the presence or absence of user contact with electronic device 500, orientation or acceleration / deceleration of electronic device 500, and temperature changes of electronic device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0094] Communication component 516 is configured to facilitate wired or wireless communication between electronic device 500 and other devices. Electronic device 500 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0095] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described multi-split air conditioning control method.
[0096] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 504 including instructions, which can be executed by a processor 520 of an electronic device 500 to complete the aforementioned multi-split air conditioning control method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0097] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described multi-split air conditioning control method when executed by the programmable device.
[0098] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.
[0099] In the above detailed description, terms such as "center," "upper," "lower," "left," and "right" indicate direction or positional relationship. Since components of the described device can be positioned in multiple different orientations, these directional terms are for illustrative purposes and not restrictive. It should be understood that other aspects can be utilized and structural or logical changes can be made without departing from the concept of this disclosure. Therefore, the following detailed description should not be considered limiting.
[0100] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other.
[0101] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the examples described herein, the first component, part, region, layer, or section mentioned in the examples may also be referred to as the second component, part, region, layer, or section. Furthermore, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include at least one of that feature. In the description herein, “a plurality” means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0102] Furthermore, the term “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term “exemplary” is intended to present the concept in a concrete manner. As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clear from the context, “X applies A or B” is intended to mean any of the natural inclusive arrangements. That is, “X applies A or B” satisfies any of the foregoing instances if X applies A; X applies B; or both X applies A and B. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles “a” and “an” as used in this application and the appended claims are generally understood to mean “one or more.”
[0103] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term “including.”
[0104] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0105] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for controlling a multi-split air conditioning system, characterized in that, The multi-split air conditioner includes one outdoor unit and multiple indoor units, and the method includes: The first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are obtained. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, indoor unit operating parameters and indoor unit environment parameters of the corresponding indoor unit. The first state space parameters of the outdoor unit and the second state space parameters of each indoor unit are processed by the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is based on the simulation environment corresponding to the multi-split air conditioner, with the minimum power consumption as the training objective, and the basic control model is trained to obtain the target control model. The outdoor unit is controlled based on the first control parameter, and the corresponding indoor unit is controlled based on the second control parameter of each indoor unit.
2. The multi-split air conditioning control method according to claim 1, characterized in that, The target control model includes multiple distributed sub-models; The process of processing the first state-space parameters of the outdoor unit and the second state-space parameters of each indoor unit using the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit includes: Determine the distributed sub-model corresponding to each external machine and each internal machine, with one external machine or one internal machine corresponding to one distributed sub-model; The first state space parameters of the outdoor unit are input into the distributed sub-model corresponding to the outdoor unit to obtain the first control parameters of the outdoor unit; The second state space parameters of each internal unit are input into the distributed sub-model corresponding to each internal unit to obtain the second control parameters of each internal unit.
3. The multi-split air conditioning control method according to claim 1, characterized in that, The target control model is trained through the following steps: Based on the multi-split air conditioner, a corresponding simulation environment is constructed, which includes a simulated outdoor unit and multiple simulated indoor units; The basic control model, including a total reward network and a control network, is trained through multiple rounds of iterative training in the simulation environment. After each round of training, determine the error loss corresponding to that round of training; Based on the aforementioned error loss, the basic control model is optimized; If the basic control model meets the preset conditions, training is stopped, and the target control model is obtained.
4. The multi-split air conditioning control method according to claim 3, characterized in that, The control network includes multiple distributed sub-networks, with one simulated outdoor unit or one simulated indoor unit corresponding to one distributed sub-network. The step of performing multiple rounds of iterative training on the basic control model in the simulation environment includes: For each round of iterative training, the first simulated state space parameters corresponding to this round of training are obtained through the simulated environment. The first simulated state space parameters include the state space parameters corresponding to the simulated outdoor unit and the state space parameters corresponding to each simulated indoor unit. The state space parameters corresponding to the simulated outdoor unit are input into the distributed sub-network corresponding to the simulated outdoor unit to obtain the first simulated reward parameter of the simulated outdoor unit, wherein the first simulated reward parameter includes the first simulated control parameter. The state space parameters corresponding to each simulated intranet are input into the distributed subnetwork corresponding to each simulated intranet to obtain the second simulated reward parameters of each simulated intranet. The second simulated reward parameters include the second simulated control parameters.
5. The multi-split air conditioning control method according to claim 4, characterized in that, The step of determining the error loss corresponding to each training round after each round includes: After each round of training, obtain the first simulated reward parameters and the second simulated reward parameters corresponding to this round of training; Input the first and second simulation control parameters corresponding to this round of training into the simulation environment to obtain the power consumption parameters corresponding to this round of training and the second simulation state space parameters corresponding to the next round of training. Based on the first simulated control parameters, the second simulated control parameters, the power consumption parameters, and the control constraints, the reward value corresponding to this round of training is obtained; Input the first simulated reward parameter, the second simulated reward parameter, and the second simulated state space parameter into the total reward network to obtain the estimated total reward for this round of training. The error loss for this training round is obtained based on the first simulated state space parameters, the second simulated state space parameters, the reward value and estimated total reward for this training round, and the estimated total reward for the previous training round.
6. The multi-split air conditioning control method according to claim 5, characterized in that, The step of obtaining the reward value corresponding to this round of training based on the first simulated control parameters, the second simulated control parameters, the power consumption parameters, and the control constraints includes: If both the first and second simulated control parameters satisfy the control constraints, the reward value corresponding to this round of training is obtained based on the power consumption parameter. If either the first simulation control parameter or the second simulation control parameter fails to meet the control constraint, the reward value corresponding to this round of training is obtained based on the maximum reward threshold.
7. The multi-split air conditioning control method according to claim 6, characterized in that, The power consumption parameters include actual power consumption, execution time, set temperature, and indoor ambient temperature. The step of obtaining the reward value corresponding to this round of training based on the power consumption parameter includes: The reward value corresponding to the simulated outdoor unit is determined based on the actual power consumption of the simulated outdoor unit. The reward value for each simulated indoor unit is determined based on its actual power consumption, execution time, set temperature, and indoor ambient temperature. The sum of the reward value corresponding to the simulated outdoor unit and the reward value corresponding to each simulated indoor unit is determined as the reward value for this round of training.
8. The multi-split air conditioning control method according to claim 7, characterized in that, The process of determining the reward value for each simulated indoor unit based on its actual power consumption, execution time, set temperature, and indoor ambient temperature includes: For any simulated indoor unit in each simulated indoor unit, if the indoor ambient temperature corresponding to the simulated indoor unit is not equal to the set temperature, the correction parameters corresponding to the simulated indoor unit are obtained based on the execution time, set temperature and indoor ambient temperature corresponding to the simulated indoor unit. The reward value corresponding to the simulated indoor unit is obtained based on the correction parameters corresponding to the simulated indoor unit and the actual power consumption. When the indoor ambient temperature corresponding to the simulated indoor unit is equal to the set temperature, the reward value corresponding to the simulated indoor unit is obtained by subtracting the actual power consumption from the maximum reward threshold.
9. The multi-split air conditioning control method according to claim 5, characterized in that, The total reward network includes a parameter generation network and a combined network; The step of inputting the first simulated reward parameter, the second simulated reward parameter, and the second simulated state space parameter into the total reward network to obtain the estimated total reward corresponding to this round of training includes: The second simulated state-space parameters are processed by the parameter generation network to obtain a first combined parameter, a second combined parameter, a first bias term, and a second bias term; Through the combined network, the first combined parameter, the first bias term, the second combined parameter, and the second bias term are sequentially combined with the first simulated reward parameter and the second simulated reward parameter to obtain the estimated total reward corresponding to this round of training.
10. The multi-split air conditioning control method according to any one of claims 1 to 9, characterized in that, The method further includes: Obtain the historical state space parameters corresponding to the multi-split air conditioner within the target time period; The target control model is updated based on the historical state space parameters to obtain the updated target control model.
11. A multi-split air conditioning control device, characterized in that, Configured in a multi-split air conditioner, the multi-split air conditioner including one outdoor unit and multiple indoor units, the device includes: The first acquisition module is configured to acquire the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit. The first state space parameters include the outdoor unit operating parameters and the outdoor unit environment parameters. The second state space parameters include the indoor unit setting parameters, the indoor unit operating parameters and the indoor unit environment parameters of the corresponding indoor unit. The first acquisition module is configured to process the first state space parameters of the outdoor unit and the second state space parameters of each indoor unit through the target control model to obtain the first control parameters of the outdoor unit and the second control parameters of each indoor unit. The target control model is obtained by training the basic control model based on the simulation environment corresponding to the multi-split air conditioner with the minimum power consumption as the training objective. The control module is configured to control the outdoor unit based on the first control parameter, and to control the corresponding indoor unit based on the second control parameter of each indoor unit.
12. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the steps of the multi-split air conditioning control method according to any one of claims 1 to 10.
13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the steps of the multi-split air conditioning control method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the multi-split air conditioning control method according to any one of claims 1 to 10.