Central air conditioning control method, system, device, medium and program product

By designing the critical network as a global and local critical network and embedding GAT in the global critical network, the conflict problem of reward functions in the central air conditioning system of commercial buildings is solved, and optimized control of efficient energy saving and thermal comfort is achieved, control performance is improved and real-time optimization is achieved.

CN120120723BActive Publication Date: 2025-08-12SHANDONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510614562.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-12
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing central air conditioning control method is difficult to effectively balance energy efficiency and thermal comfort in commercial buildings. The conflict between reward functions in deep reinforcement learning of multiple agents leads to unstable learning strategies and fails to effectively model the relationship between system energy consumption coupling and thermal coupling, resulting in a decline in control performance.

Method used

The multi-agent deep reinforcement learning method that decomposes the critic network and the graph attention network is adopted, and the critical network is designed as a global critical network and a local critical network. The graph attention neural network (GAT) is embedded to alleviate reward conflicts, and the coupled information within the system is fully extracted to build an optimized control model for the actor network, global critical network and local critical network.

Benefits of technology

It realizes efficient energy saving and thermal comfort control of central air conditioning systems in commercial buildings, eliminates unstable strategies, improves optimized control performance, and realizes real-time optimized control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120120723B_ABST
    Figure CN120120723B_ABST
Patent Text Reader

Abstract

The present invention discloses a central air conditioning control method, system, device, medium, and program product, relating to the technical field of data processing and control systems. The method comprises constructing a central air conditioning system control model; constructing an optimization control model consisting of an actor network, a global critic network embedded with a GAT, and a local critic network; and training the optimization control model, whereby the trained actor network controls the cooling and heating temperature setpoints for each region based on the region's current state observation variables. The critic network is designed as a global critic network and a local critic network to alleviate the conflict between regional thermal comfort rewards and central air conditioning system total energy consumption rewards during training. A graph attention network is integrated into the global critic network to fully extract coupling information within the system and improve optimization control performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing and control systems, and in particular to a central air-conditioning control method, system, equipment, medium and program product. Background Art

[0002] Existing central air conditioning control methods are mainly divided into three categories:

[0003] (1) Rule-based controls (RBC) method: This method relies on fixed rules and lacks flexibility in nature. Therefore, it is difficult to effectively balance energy efficiency and thermal comfort in central air-conditioning systems of complex commercial buildings.

[0004] (2) Model Predictive Control (MPC) method: This method optimizes energy consumption and thermal comfort through dynamic models, but it is highly dependent on the accuracy of the dynamic models. However, due to the uncertainty of building thermal parameters and the complexity of the central air-conditioning system operation mechanism, it is quite challenging to develop an accurate dynamic model for the central air-conditioning system of commercial buildings.

[0005] (3) Deep Reinforcement Learning (DRL) method, which learns the optimal control strategy through direct interaction with the environment and does not require complex dynamic models, making it suitable for HVAC (Heating Ventilation and Air Conditioning) systems.

[0006] Deep reinforcement learning is divided into single-agent deep reinforcement learning and multi-agent deep reinforcement learning (MADRL). Due to the interactions between various areas in the central air-conditioning system of commercial buildings, single-agent reinforcement learning cannot effectively achieve collaborative optimization among various agents.

[0007] Although MADRL can effectively coordinate collaborative optimization between regions, the significant energy consumption coupling and thermal coupling characteristics of commercial building central air-conditioning systems pose certain challenges to MADRL. First, the optimization objectives of commercial building central air-conditioning systems are the total energy consumption of the central air-conditioning system and the thermal comfort of each region. Therefore, MADRL's reward function is usually a combination of regional thermal comfort rewards and total energy consumption rewards for the central air-conditioning system. However, this reward combination will conflict with each other during training, resulting in unstable learning strategies. In addition, MADRL focuses on policy coordination between agents, but does not effectively model the relationship between system energy consumption coupling and thermal coupling, resulting in reduced control performance. Summary of the Invention

[0008] To address the above-mentioned issues, the present invention proposes a central air-conditioning control method, system, device, medium, and program product. The critic network in multi-agent reinforcement learning is designed as a global critic network and a local critic network. This effectively alleviates the conflict between regional thermal comfort rewards and total energy consumption rewards of the central air-conditioning system during training, eliminating unstable strategies. At the same time, a graph attention network (GAT) is integrated into the global critic network to fully extract coupling information within the system, thereby improving optimization control performance.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] In a first aspect, the present invention provides a central air conditioning control method, comprising:

[0011] Construct a central air conditioning system control model for N zones, including state observation variables defined by the indoor and outdoor temperature and humidity of each zone, actions defined by the cooling and heating temperature setpoints of each zone, and rewards defined by the total energy consumption of the central air conditioning system and the thermal comfort of each zone.

[0012] An optimization control model is constructed, consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs control actions for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward.

[0013] The optimization control model is trained, and the trained actor network is used to control the cooling and heating temperature set values of each area of the central air-conditioning system according to the current state observation variables of each area.

[0014] As an optional implementation, the state observation variables include outdoor air temperature, outdoor air relative humidity, number of occupants in each area, indoor air temperature, indoor air relative humidity, heating temperature set value, cooling temperature set value and total energy consumption of the central air conditioning system;

[0015] The action includes incremental adjustment of the heating temperature set point and the cooling temperature set point, setting a range of the incremental adjustment, and a range of values of the heating temperature set point and the cooling temperature set point, and the heating temperature set point is lower than the cooling temperature set point;

[0016] Rewards include global rewards and local rewards ;

[0017] ;

[0018] ; ;

[0019] in, and is an adjustable weight; is the total energy consumption of the central air-conditioning system; is the thermal comfort bonus of region i at time t, is the occupancy status of area i at time t, where unoccupied is 0 and occupied is 1; is the thermal comfort threshold; is the predicted dissatisfaction percentage of region i at time t.

[0020] As an optional implementation, the global critic network embedded in GAT includes a GAT layer and three fully connected layers after the GAT layer. ;in, For the GAT layer, It is a three-layer fully connected layer; is the state observation variable of all regions; For all areas of action; are the parameters of the global critic network embedded in GAT.

[0021] As an optional implementation, the processing of the global critic network embedded in the GAT includes:

[0022] Each region corresponds to an agent. In the GAT layer, each agent is a node in the graph. The feature vector of each node is composed of the state observation variable-action pair of the corresponding agent. The input of the GAT layer is the feature vector of all nodes. The attention coefficient of each node is calculated. :

[0023] ;

[0024] Use the attention coefficient to aggregate the neighboring nodes of node i to update the feature vector of node i;

[0025] ;

[0026] in, is a learnable weight matrix, is a learnable attention weight vector, || is a vector concatenation operation, represents transposition, LeakyReLU(.) is the activation function, is the set of neighbor nodes of node i; is the feature vector of node i, node j, and node k; σ(.) is the activation function, is the updated feature vector of node i; It is the state dimension after node aggregation;

[0027] The updated feature vectors of all nodes are concatenated to obtain a global feature vector, which is then passed through three fully connected layers to generate a global state-action value.

[0028] As an optional implementation, during training, the loss function for updating the parameters of the global critic network embedded in GAT is for:

[0029] ;

[0030] Among them, the target value for:

[0031] ;

[0032] Where γ is the discount factor, is the global critic network embedded in GAT for region i The target network, The parameters are , The update is updated using the soft update mechanism. The parameters are , is the target network of the actor network in region i, The parameters are , min is the minimum value operation; is the state observation variable of all regions; For all areas of action; A global reward for all regions; is the global reward of region i; is the global state observation variable at the next sampling moment; D is the experience replay buffer pool; is the state observation variable of the sampling area i at the next moment; E is the expected operation function; is the action computed by the target actor network for region i.

[0033] As an alternative implementation, the loss function for updating the local critic network parameters is for:

[0034] ;

[0035] Among them, the target value for:

[0036] ;

[0037] in, For the local critic network The target network, The parameters are , The update is updated using the soft update mechanism. The parameters are ; Local rewards for all regions; is the state observation variable of region i; is the action of region i; is the local reward of region i.

[0038] As an alternative implementation, through the policy gradient function Update the parameters of the actor network;

[0039] ;

[0040] The first part on the right side of the equation is the global policy gradient calculated by the global critic network embedded in GAT, and the second part is the local policy gradient calculated by the local critic network; is the state observation variable of all regions; is the action of all areas; D is the experience replay buffer pool; For actor networks Parameters, The update is updated using the soft update mechanism; represents the gradient; is the state observation variable of region i; is the action of region i; is the first global critic network embedded in GAT, N is the total number of regions; is a local critic network.

[0041] In a second aspect, the present invention provides a central air-conditioning control system, comprising:

[0042] a control model construction module configured to construct a central air-conditioning system control model including N zones, including state observation variables defined by indoor and outdoor temperature and humidity of each zone, actions defined by cooling and heating temperature setpoints of each zone, and rewards defined by total energy consumption of the central air-conditioning system and thermal comfort of each zone;

[0043] An optimization control model construction module is configured to construct an optimization control model consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs a control action for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward.

[0044] The control module is configured to train the optimization control model, and use the trained actor network to control the cooling temperature set value and the heating temperature set value of each zone of the central air-conditioning system according to the current state observation variables of each zone.

[0045] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0047] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which implements the method described in the first aspect when executed by a processor.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] This paper proposes a central air conditioning control method, specifically a multi-agent deep reinforcement learning central air conditioning control method based on a decomposition critic network and a graph attention network, for achieving efficient energy saving and thermal comfort control of central air conditioning systems in commercial buildings. It has the following advantages:

[0050] (1) Efficiency: By designing the critic network in multi-agent reinforcement learning into a global critic network and a local critic network, the conflict between the regional thermal comfort reward and the total energy consumption reward of the central air-conditioning system during training can be effectively alleviated, eliminating unstable strategies. At the same time, by integrating the graph attention network into the global critic network, the coupling information within the system can be fully extracted, thereby improving the optimization control performance.

[0051] (2) Energy saving: It can realize energy-saving control of central air-conditioning system while ensuring thermal comfort of users of central air-conditioning system in commercial buildings.

[0052] (3) Real-time performance: Simply deploy the trained actor network in the central air-conditioning system of a commercial building. Since the forward propagation speed of the actor network is extremely fast, real-time optimization control of the central air-conditioning system can be achieved.

[0053] The present invention proposes a central air-conditioning control method, system, device, medium and program product. First, the optimization control problem of the central air-conditioning system of a commercial building is modeled as a Markov game, and the corresponding state, action and reward are constructed. Secondly, an actor and critic network architecture for multi-agent reinforcement learning is designed, wherein the actor network takes the state observation of a single agent as input and outputs the control action for the central air-conditioning system. In order to resolve the conflict between the regional thermal comfort reward and the total energy consumption reward of the central air-conditioning system during training, the critic network is designed into a global critic network and a local critic network. The global critic network takes the state observations and actions of all agents as input and outputs the global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single agent as input and outputs the local state-action value corresponding to the regional thermal comfort reward. At the same time, in order to fully model and learn the energy consumption coupling and thermal coupling relationship within the central air-conditioning system of a commercial building, a GAT is embedded in the global critic network to improve the optimization control performance.

[0054] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0056] Figure 1 This is a flow chart of the central air conditioning control method provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0057] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0058] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0059] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "comprise" and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0060] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0061] Example 1

[0062] This embodiment proposes a central air conditioning control method, specifically a multi-agent deep reinforcement learning central air conditioning control method based on decomposition critic network and graph attention network, such as Figure 1 Shown, including:

[0063] Construct a central air conditioning system control model for N zones, including state observation variables defined by the indoor and outdoor temperature and humidity of each zone, actions defined by the cooling and heating temperature setpoints of each zone, and rewards defined by the total energy consumption of the central air conditioning system and the thermal comfort of each zone.

[0064] An optimization control model is constructed, consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs control actions for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward.

[0065] The optimization control model is trained, and the trained actor network is used to control the cooling and heating temperature set values of each area of the central air-conditioning system according to the current state observation variables of each area.

[0066] In this example, the method targets a typical central air conditioning system in a commercial building with N zones. The optimization control problem for this system is modeled as a Markov game (MG) involving N agents. The state observation variables are defined by the indoor and outdoor temperature and humidity of each zone (agent), the actions are defined by the cooling and heating setpoints for each zone (agent), and the rewards are defined by the total energy consumption of the central air conditioning system and the thermal comfort of each zone.

[0067] In this embodiment, the process of constructing the MG is specifically as follows:

[0068] The optimal control problem of a central air-conditioning system in a commercial building with N zones is modeled as a MG consisting of N agents, where each zone i corresponds to an agent i.

[0069] MG consists of the following components: a set of global state sets , the action set of each agent , and the set of local state observations of each agent Each agent i uses a random strategy To select an action, the policy is based on the state transition function Generate the next global state, each agent i also has a reward .

[0070] In this embodiment, the reward is divided into two parts: global and local. Determined by the global state and actions of all agents, local rewards It only depends on the local state observation and action of agent i. Since MADRL does not require information about the state transition function, this embodiment mainly focuses on designing the state, action and reward of MG.

[0071] MG's status, actions, and rewards are as follows:

[0072] (1) Status:

[0073] For region i ( ), the corresponding local observation of agent i at time t Defined as:

[0074] (1);

[0075] Where N is the total number of regions; is the index of the day of the week, is the index of the hour of the day, is the time of day index (its value depends on the selected interval, for example, if the interval is 15 minutes, then ); is the outdoor air temperature; is the relative humidity of outdoor air; is the number of people occupying area i; is the indoor air temperature; is the relative humidity of indoor air; is the heating temperature setting value; is the cooling temperature set point; is the total energy consumption of the central air-conditioning system.

[0076] Finally, the local state observations of all agents are expressed as , this embodiment selects As global state .

[0077] (2) Action:

[0078] Agent i at time Action Defined as:

[0079] (2);

[0080] in, For incremental adjustments to the heating temperature setpoint, For incremental adjustment of the cooling temperature set point, the incremental adjustment range is limited to Inside.

[0081] In addition, the heating temperature setting value and cooling temperature setpoint Restricted to The heating temperature setting value is always lower than the cooling temperature setting value. Finally, the actions of all agents are expressed as .

[0082] (3) Rewards:

[0083] The role of rewards is to provide immediate feedback for the actions of the agent. For the optimization control problem of the central air-conditioning system in a commercial building, the reward function must balance two key indicators: energy consumption and thermal comfort. This embodiment uses the total energy consumption of the central air-conditioning system As the energy consumption index, the predicted percentage of dissatisfied (PPD) is used as the thermal comfort evaluation index.

[0084] In order to eliminate unit differences, the total energy consumption and PPD of the central air-conditioning system are normalized. At the moment Rewards Defined as:

[0085] (3);

[0086] in, is the normalized PPD; and are adjustable weights used to balance energy consumption and thermal comfort indicators according to the optimization goal; Indicates the occupancy status of area i at time t, where unoccupied is 0 and occupied is 1.

[0087] To ensure that the PPD value remains within the thermal comfort threshold of the ASHRAE standard (ASHRAE standard is the abbreviation of the American Society of Heating, Refrigerating and Air-Conditioning Engineers, Inc., translated as the American Society of Heating, Refrigerating and Air-Conditioning Engineers, Inc.), the normalized thermal comfort bonus Redefine as:

[0088] (4);

[0089] in, is the thermal comfort threshold; is the PPD of region i at time t.

[0090] During MADRL training, the PPD reward focuses on the thermal comfort of a single area, while the total energy consumption reward of the central air conditioning system reflects the collective behavior of all areas. This entanglement between local rewards and global rewards may lead to unstable strategies. To alleviate this problem, this embodiment will reward Decomposed into global rewards and local rewards ;

[0091] (5);

[0092] (6);

[0093] Finally, the global rewards and local rewards of all agents are expressed as and .

[0094] In this embodiment, an actor and critic network architecture for multi-agent reinforcement learning is constructed. The actor network takes the state observations of regional agents as input and outputs the optimal control action for the central air-conditioning system. At the same time, in order to solve the entanglement problem between the regional thermal comfort reward and the total energy consumption reward of the central air-conditioning system during training, the critic network is designed into a global critic network and a local critic network. The global critic network takes the state observations and actions of all regional agents as input and outputs a global state-action value corresponding to the total energy consumption of the central air-conditioning system. The local critic network takes the state observations and actions of each regional agent as input and outputs a local state-action value related to the evaluation of regional thermal comfort. In addition, in order to fully model and learn the energy consumption coupling and thermal coupling within the central air-conditioning system of a commercial building, a GAT is embedded in the global critic network to improve the optimization control performance.

[0095] Specifically:

[0096] Each agent has two GAT-based global critic networks, a local critic network, an actor network and its corresponding target network.

[0097] Specifically, for agent i, the actor network is represented as , the parameters are , input local state observation , output action , output action As input to the local critic network and the global critic network, perform the output action The final state observation serves as the input to the local critic network and the global critic network.

[0098] The local critic network of agent i is represented as , the parameters are , observing the local state of agent i and actions As input, the local state-action value corresponding to the local reward (regional thermal comfort reward) is calculated; the local state observation input to the local critic network is sampled from the experience replay pool, and the set of data in the experience replay pool is composed of the state observations obtained after executing the actions output by the actor network.

[0099] Two GAT-based global critic networks for agent i and The parameters are and , and then take the minimum value of the output of these two networks during training, as shown in formula (15); these two networks observe the state of all agents and actions As input, the global state-action value corresponding to the global reward (the total energy consumption reward of the central air-conditioning system) is calculated.

[0100] Specifically, in order to fully model and learn the energy coupling and thermal coupling relationships within the central air-conditioning system of commercial buildings, the GAT-based global critic network consists of one GAT layer and three fully connected layers.

[0101] For the agent , its global critic network is as follows (7):

[0102] (7);

[0103] in, For the GAT layer, represents the three fully connected layers applied after the GAT layer; A global critic network embedded in GAT Parameters.

[0104] In the GAT layer, each agent is considered as a node in the graph, and its feature vector consists of the state-action pair of the corresponding agent. The input of the GAT layer includes the feature vectors of all nodes, which is expressed as Equation (8):

[0105] (8);

[0106] in, is the eigenvector of node i, is the dimension of the feature vector, and dim(.) is the operation to obtain the vector dimension.

[0107] GAT uses the attention mechanism to assign different importance to the neighboring nodes of each node, thereby modeling and learning the interaction between agents. The attention coefficient of node i is Calculation is as follows (9):

[0108] (9);

[0109] in, is a learnable weight matrix used to transform node feature vectors, is a learnable attention weight vector, || represents a vector concatenation operation, represents transposition, LeakyReLU(.) is the activation function, represents the set of neighbor nodes of node i; is the feature vector of node i, node j, and node k; It is the state dimension after node aggregation.

[0110] Afterwards, the attention coefficient is used to aggregate the neighbor node information of node i to update the feature vector of node i, which is calculated as formula (10):

[0111] (10);

[0112] Among them, σ(.) is the activation function, j represents the jth neighbor node, is the eigenvector of node j, is the updated feature vector of node i.

[0113] After updating the feature vectors of all nodes, concatenate them into a global feature vector H:

[0114] (11).

[0115] Finally, the global feature vector is passed through three fully connected layers to generate the final global state-action value:

[0116] (12).

[0117] As you can understand, the actor network takes the observation variables of a single agent as input and outputs the optimal control action for the central air-conditioning system, and the local critic network takes the observation variables and actions of a single agent as input and outputs the local state-action value corresponding to the regional thermal comfort reward. The processing process of the two networks belongs to the regular processing flow of the actor network and the critic network, and will not be repeated here.

[0118] In this embodiment, the intelligent agent interacts extensively with a commercial building central air conditioning system simulation environment to collect training data and train the actor network, global critic network, and local critic network.

[0119] moments during training , agent i obtains local state observation , and according to its actor network Generate Action , as follows:

[0120] (13);

[0121] in, for Gaussian noise at the moment.

[0122] Subsequently, agent i receives the global reward and local rewards , and obtain the local state observation at the next moment When all N agents complete this cycle, the multi-agent system will send the experience tuple Stored in the experience replay buffer pool D.

[0123] When the number of experiences in the experience replay buffer pool D reaches a threshold, a mini-batch is randomly sampled from D. To train the actor network and the critic network; where m is the index subscript of the experience data; K is the number of experiences taken from the experience revisit pool each time.

[0124] Specifically, using the loss function shown in formula (14) Update the parameters of the global critic network:

[0125] (14).

[0126] Among them, the target value Defined as:

[0127] (15);

[0128] Where γ is the discount factor, is the global critic network embedded in GAT for agent i The target network, The parameters are , The parameters are , is the target network of the actor network of agent i, The parameters are ; Represents sampling from the experience replay buffer pool D; is the state observation variable of all regions; For all areas of action; is the global reward for all sampled areas; is the global reward of agent i; is the global state observation variable obtained by sampling The corresponding global state observation variable at the next moment; D is the experience replay buffer pool; is the state observation variable of the sampled agent i at the next moment; min is the minimum operation. Applying this operation to the target network of the two global critic networks can effectively alleviate the overestimation problem of Q value; E is the expected operation function; is the action computed by the target actor network for agent i.

[0129] By implementing a delayed policy update mechanism, the actor network's dependence on the critic network is significantly reduced. To this end, this embodiment chooses to delay the update of the local critic network and actor network every T steps.

[0130] Similarly, based on the loss function shown in formula (16) Update the parameters of the local critic network of agent i:

[0131] (16).

[0132] Among them, the target value Defined as:

[0133] (17);

[0134] in, For the local critic network The target network, The parameters are , The parameters are ; Local rewards for all regions; is the state observation variable of region i; is the action of region i; is the local reward of region i.

[0135] Through the policy gradient function shown in formula (18) Update the parameters of the actor network of agent i.

[0136] (18);

[0137] in, is the state observation variable of all regions; is the action of all areas; D is the experience replay buffer pool; For actor networks Parameters; represents the gradient; is the state observation variable of region i; is the action of region i; is the first global critic network embedded in GAT, N is the total number of regions; is a local critic network.

[0138] Equation (18) consists of two parts: the first is the global policy gradient calculated by the global critic network, and the second is the local policy gradient calculated by the local critic network. This decomposed policy gradient can effectively alleviate the problem of policy instability caused by the entanglement of global and local reward signals.

[0139] In addition, each target network of agent i is updated through the following soft update mechanism to obtain the parameters of each updated target network: 、 、 :

[0140] (19);

[0141] (20);

[0142] (twenty one);

[0143] in, is the soft update coefficient, which is used to adjust the update rate of the target network.

[0144] In this embodiment, after training is completed, the actor network model parameters are fixed, and the actor network is deployed in the central air-conditioning system of a commercial building to achieve real-time optimization control.

[0145] In this embodiment, the optimization control problem of the central air-conditioning system of a commercial building is first modeled as a Markov game, and the corresponding states, actions, and rewards are constructed. Secondly, an actor and critic network architecture for multi-agent reinforcement learning is designed, in which the actor network takes the state observation of a single agent as input and outputs the control action for the central air-conditioning system. In order to resolve the conflict between the regional thermal comfort reward and the total energy consumption reward of the central air-conditioning system during training, the critic network is designed as a global critic network and a local critic network. The global critic network takes the state observations and actions of all agents as input and outputs the global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single agent as input and outputs the local state-action value corresponding to the regional thermal comfort reward. At the same time, in order to fully model and learn the energy consumption coupling and thermal coupling relationship within the central air-conditioning system of a commercial building, a GAT is embedded in the global critic network to improve the optimization control performance. Finally, through extensive interaction with the simulation environment of the commercial building's central air-conditioning system, the actor network, global and local critic networks, and their target networks are trained, and the trained actor network is deployed to achieve real-time optimization control of the commercial building's central air-conditioning system.

[0146] Example 2

[0147] This embodiment provides a central air conditioning control system, including:

[0148] a control model construction module configured to construct a central air-conditioning system control model including N zones, including state observation variables defined by indoor and outdoor temperature and humidity of each zone, actions defined by cooling and heating temperature setpoints of each zone, and rewards defined by total energy consumption of the central air-conditioning system and thermal comfort of each zone;

[0149] An optimization control model construction module is configured to construct an optimization control model consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs a control action for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward.

[0150] The control module is configured to train the optimization control model, and use the trained actor network to control the cooling temperature set value and the heating temperature set value of each zone of the central air-conditioning system according to the current state observation variables of each zone.

[0151] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0152] In further embodiments, there is also provided:

[0153] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed by the processor, wherein when the computer instructions are executed by the processor, the method described in Example 1 is performed. For the sake of brevity, no further details are given here.

[0154] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0155] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0156] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in Example 1 is performed.

[0157] The method in Example 1 can be directly implemented as a hardware processor, or can be implemented using a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, it will not be described in detail here.

[0158] A computer program product includes a computer program, which implements the method described in embodiment 1 when executed by a processor.

[0159] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions contained in program modules, which are executed in a device on a real or virtual processor of a target to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided between program modules as needed. The machine-executable instructions for the program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0160] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0161] In the context of the present invention, computer program code or related data can be carried by any appropriate carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, and the like.

[0162] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0163] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A central air conditioning control method, characterized in that: include: Construct a central air conditioning system control model for N zones, including state observation variables defined by the indoor and outdoor temperature and humidity of each zone, actions defined by the cooling and heating temperature setpoints of each zone, and rewards defined by the total energy consumption of the central air conditioning system and the thermal comfort of each zone. An optimization control model is constructed, consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs control actions for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward. The optimization control model is trained, and the trained actor network is used to control the cooling and heating temperature set points of each zone in the central air-conditioning system according to the current state observation variables of each zone. The state observation variables include outdoor air temperature, outdoor air relative humidity, number of occupants in each area, indoor air temperature, indoor air relative humidity, heating temperature setpoint, cooling temperature setpoint, and total energy consumption of the central air conditioning system; The action includes incremental adjustment of the heating temperature set point and the cooling temperature set point, setting a range of the incremental adjustment, and a range of values of the heating temperature set point and the cooling temperature set point, and the heating temperature set point is lower than the cooling temperature set point; Rewards include global rewards and local rewards ; ; ; ; in, and is an adjustable weight; is the total energy consumption of the central air-conditioning system; is the thermal comfort bonus of region i at time t, is the occupancy status of area i at time t, where unoccupied is 0 and occupied is 1; is the thermal comfort threshold; is the predicted dissatisfaction percentage of region i at time t; The global critic network embedded in GAT consists of a GAT layer and three fully connected layers after the GAT layer. ;in, For the GAT layer, It is a three-layer fully connected layer; is the state observation variable of all regions; For all areas of action; A global critic network embedded in GAT Parameters; During training, the loss function for updating the parameters of the global critic network embedded in GAT is for: ; Among them, the target value for: ; Where γ is the discount factor, is the global critic network embedded in GAT for region i The target network, The parameters are , The update is updated using the soft update mechanism. The parameters are , is the target network of the actor network in region i, The parameters are , min is the minimum value operation; is the state observation variable of all regions; For all areas of action; A global reward for all regions; is the global reward of region i; is the global state observation variable at the next sampling moment; D is the experience replay buffer pool; is the state observation variable of the sampling area i at the next moment; E is the expected operation function; is the action calculated by the target actor network in region i; Through the policy gradient function Update the parameters of the actor network; ; The first part on the right side of the equation is the global policy gradient calculated by the global critic network embedded in GAT, and the second part is the local policy gradient calculated by the local critic network; is the state observation variable of all regions; is the action of all areas; D is the experience replay buffer pool; For actor networks Parameters, The update is updated using the soft update mechanism; represents the gradient; is the state observation variable of region i; is the action of region i; is the first global critic network embedded in GAT, N is the total number of regions; is a local critic network.

2. A central air conditioning control method according to claim 1, characterized in that: The processing of the global critic network embedded in GAT includes: Each region corresponds to an agent. In the GAT layer, each agent is a node in the graph. The feature vector of each node is composed of the state observation variable-action pair of the corresponding agent. The input of the GAT layer is the feature vector of all nodes. The attention coefficient of each node is calculated. : ; Use the attention coefficient to aggregate the neighboring nodes of node i to update the feature vector of node i; ; in, is a learnable weight matrix, is a learnable attention weight vector, || is a vector concatenation operation, represents transposition, LeakyReLU(.) is the activation function, is the set of neighbor nodes of node i; is the feature vector of node i, node j, and node k; σ(.) is the activation function, is the updated feature vector of node i; It is the state dimension after node aggregation; The updated feature vectors of all nodes are concatenated to obtain a global feature vector, which is then passed through three fully connected layers to generate a global state-action value.

3. A central air conditioning control method according to claim 1, characterized in that: Loss function for updating local critic network parameters for: ; Among them, the target value for: ; in, For the local critic network The target network, The parameters are , The update is updated using the soft update mechanism. The parameters are ; Local rewards for all regions; is the state observation variable of region i; is the action of region i; is the local reward of region i.

4. A central air conditioning control system, using a central air conditioning control method according to any one of claims 1 to 3, characterized in that: include: a control model construction module configured to construct a central air-conditioning system control model including N zones, including state observation variables defined by indoor and outdoor temperature and humidity of each zone, actions defined by cooling and heating temperature setpoints of each zone, and rewards defined by total energy consumption of the central air-conditioning system and thermal comfort of each zone; An optimization control model construction module is configured to construct an optimization control model consisting of an actor network, a global critic network embedded in a GAT, and a local critic network. The actor network takes the state observation variables of a single region as input and outputs a control action for the central air-conditioning system. The global critic network embedded in the GAT takes the state observation variables and actions of all regions as input and outputs a global state-action value corresponding to the total energy consumption reward of the central air-conditioning system. The local critic network takes the state observation variables and actions of a single region as input and outputs a local state-action value corresponding to the regional thermal comfort reward. The control module is configured to train the optimization control model, and use the trained actor network to control the cooling temperature set value and the heating temperature set value of each zone of the central air-conditioning system according to the current state observation variables of each zone.

5. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 3 is completed.

6. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, complete the method according to any one of claims 1 to 3.

7. A computer program product, characterized in that The invention comprises a computer program, which is used to implement the method according to any one of claims 1 to 3 when executed by a processor.

Citation Information

Patent Citations

  • Unmanned aerial vehicle group cooperative combat method based on multi-agent reinforcement learning

    CN117055623A

  • Calculation unloading and resource allocation method based on GAT mixed action multi-agent reinforcement learning

    CN117098189A

  • Building air conditioner optimization control method and system considering personalized comfort of regional users

    CN118794109A