Deep reinforcement learning method and system for building energy control
By constructing a deep reinforcement learning method with dual experience pools and using a large language model to correct the experience data in building energy control, the problems of low efficiency and easy falling into suboptimal strategies in existing technologies are solved, and efficient and stable building energy control is achieved.
Patent Information
- Application Number
- CN202511036785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing deep reinforcement learning methods are inefficient in building energy control and are prone to falling into suboptimal strategies, resulting in low control accuracy, response delays and system instability, affecting energy utilization efficiency and grid burden.
A dual experience pool based on a large language model is combined with a deep reinforcement learning algorithm. By constructing the state space, action space and reward function, the large language model is used to correct the experience data and build an intelligent agent to quickly explore the optimal control strategy and reduce invalid exploration.
It improves the learning efficiency and overall performance of building energy control strategies, ensures that intelligent agents can quickly find the optimal control strategy in complex environments, and improves the stability and energy utilization efficiency of the system.
Smart Images

Figure CN120542515B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of deep reinforcement learning technology, and in particular relates to a deep reinforcement learning method and system for building energy control. Background Art
[0002] Buildings account for approximately 40% of global energy consumption and carbon emissions, making them a critical sector in the global transition to a low-carbon future. Efficient and adaptive building energy control is a major challenge, balancing multiple objectives such as user comfort, grid reliability, and energy efficiency. Deep reinforcement learning offers a promising solution, enabling intelligent agents to learn control policies through interaction with their environment and adapt to dynamic changes using reward feedback, balancing multiple control objectives.
[0003] However, due to the large number of uncertain system parameters in building energy systems, the state space is vast and the optimal control strategies are limited. When faced with such complex decision-making challenges, traditional deep reinforcement learning methods often face the dual challenges of accuracy and efficiency. Existing deep reinforcement learning methods rely on extensive trial-and-error random exploration, which easily produces a large number of low-quality experience samples. These low-quality experience samples hinder policy optimization, affecting learning efficiency and potentially leading to poor control strategies. This leads to low control accuracy, delayed response, and compromised system stability, ultimately resulting in energy waste, increased grid burden, and increased security risks. Summary of the Invention
[0004] The embodiments of the present application provide a deep reinforcement learning method and device for building energy control, which can solve the problems of low sampling efficiency and low learning efficiency of existing deep reinforcement learning methods and the susceptibility to falling into suboptimal strategies.
[0005] In a first aspect, embodiments of the present application provide a deep reinforcement learning method for building energy control, comprising:
[0006] An intelligent agent is constructed based on a target building environment; the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control action performed on the target building environment, and the reward function represents the reward value obtained after performing the control action under the environmental state;
[0007] collecting common experience data each time the intelligent agent interacts with the target building environment, and storing the common experience data in a common experience pool;
[0008] Constructing a prompt text according to the state space, the action space and the intention description text; the intention description text is used to indicate the action range of the control action that can be executed under each of the environmental states;
[0009] Inputting the prompt text into a large language model to obtain an action range of each of the control actions in each of the environmental states;
[0010] Correcting the common experience data according to the motion range to obtain LLM-corrected experience data, and storing the LLM-corrected experience data in an LLM experience pool;
[0011] The intelligent agent is trained according to the common experience pool and the LLM experience pool by minimizing the loss function to obtain a trained building energy control strategy model.
[0012] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0013] The system leverages the prior knowledge embedded in the large language model to analyze various environmental states and infers an adaptive action range based on control intent. This guides and corrects low-quality experience data during the agent's random exploration, reducing ineffective exploration. By constructing a dual experience pool that stores both random and high-quality experience data, combined with a deep reinforcement learning algorithm, the agent can quickly and effectively explore the optimal control strategy, improving the efficiency and overall performance of building energy control strategy learning.
[0014] In a possible implementation of the first aspect, the step of constructing an intelligent agent based on the target building environment includes:
[0015] The state space is constructed according to the detection data of the target building environment. The expression of the state space is:
[0016] ,
[0017] in, represents the environmental state at time t, Indicates the calendar class status, Indicates weather conditions. represents the community-level status, Indicates building-level status;
[0018] The action space is constructed according to the control behavior data of the building energy system of the target building environment. The expression of the action space is:
[0019] ,
[0020] in, represents the control action executed at time t, represents the charge and discharge ratio of the hot water energy storage system, Expressed as the charge and discharge ratio of battery energy storage, Expressed as the output power of the heat pump;
[0021] Taking any environmental state in the state space as the current environmental state, the state transition probability is determined based on any control action in the action space. The expression of the state transition probability is:
[0022] ,
[0023] in, represents the state transition probability, represents the environmental state at time t+1;
[0024] The reward function is constructed based on the immediate reward and discount factor for executing any control action in the action space under the current environment state. The expression of the reward function is:
[0025] ,
[0026] in, Indicates cumulative rewards, represents the number of training steps, represents the discount factor that balances current rewards and future rewards, Indicates immediate reward;
[0027] The intelligent agent is constructed based on the state space, the action space, the state transition probability and the reward function.
[0028] In the above scheme, the building energy control problem is modeled as a Markov decision process, and the intelligent agent is designed according to the state space, action space and reward function in the Markov decision. This does not require any prior information about uncertain parameters and a clear building thermal dynamics model, and has wider applicability.
[0029] In a possible implementation of the first aspect, the deep reinforcement learning method for building energy control further includes constructing the immediate reward, and the step of constructing the immediate reward includes:
[0030] A comfort reward is constructed according to the indoor temperature of the target building environment. The expression of the comfort reward is:
[0031] ,
[0032] in, Indicates comfort bonus, represents the indoor temperature at time t, Indicates the set comfort temperature. Indicates the set temperature deviation threshold;
[0033] A grid reward is constructed according to the grid volatility and grid power consumption of the target building environment. The expression of the grid reward is:
[0034] ,
[0035] in, represents the grid reward, Indicates the grid volatility, The numerical range of ), represents the symbolic function, represents the power consumption of the grid at time t, represents the power consumption of the grid at time t-1, and represents the weight coefficient;
[0036] The instant reward is constructed according to the comfort reward and the grid reward. The expression of the instant reward is:
[0037] ,
[0038] in, and Represents the weight coefficient.
[0039] In the above scheme, multi-objective weighted rewards are used to decompose complex task objectives into multiple quantifiable objectives of comfort and energy consumption, and through weight adjustment, the intelligent agent is guided to find a balance between multiple objectives, avoiding a single objective dominating the strategy and ensuring the optimal overall performance of the strategy.
[0040] In a possible implementation of the first aspect, the step of collecting common experience data each time the intelligent agent interacts with the target building environment includes:
[0041] The agent interacts with the target building environment at a current moment to obtain a current environment state, and uses the agent to select and execute a first control action to obtain a next environment state returned by the target building environment and a reward value obtained by the agent for executing the first control action;
[0042] The general experience data is constructed based on the current environment state, the first control action, the next environment state, and the reward value.
[0043] In a possible implementation of the first aspect, the step of correcting the general experience data according to the motion range to obtain experience data corrected based on the LLM includes:
[0044] Acquiring a target action range of the first control action;
[0045] Correcting the first control action to a second control action that meets the target action range;
[0046] The LLM-corrected experience data is constructed based on the current environment state, the second control action, the next environment state, and the reward value.
[0047] In a possible implementation of the first aspect, the intelligent agent corresponds to a policy network and a value network, and the step of training the intelligent agent by minimizing a loss function based on the normal experience pool and the LLM experience pool includes:
[0048] Random sampling is performed from the common experience pool and the LLM experience pool according to a preset sampling ratio to obtain a batch of sampling experience data;
[0049] The standard policy network loss is calculated based on the sampled experience data and the policy network loss function. The expression of the standard policy network loss is:
[0050] ,
[0051] in, represents the standard strategy network loss, represents the network parameters of the policy network, represents the computational expectation, Indicates from Each piece of sampled experience data, Indicates the experience data of the ordinary experience pool and the experience data of the LLM experience pool, represents the entropy coefficient of the policy network, represents the strategy generated by the policy network, represents the predicted Q value of the value network, Network parameters representing the value network;
[0052] Updating the network parameters of the policy network by minimizing the standard policy network loss;
[0053] The value network loss is calculated based on the sampled empirical data and the value network loss function. The expression of the value network loss is:
[0054] ,
[0055] in, represents the value network loss, represents the target Q value;
[0056] The network parameters of the value network are updated by minimizing the value network loss.
[0057] In this approach, efficient learning is achieved through two collaborative networks. The standard policy network loss guides the policy network to optimize action selection, while the value network loss guides the value network to accurately estimate action values. Together, these two networks reduce gradient variance and improve learning stability. Furthermore, experience data is sampled from two separate experience pools and randomly shuffled, ensuring both diversity and high-value experience data for the reinforcement learning agent. This promotes effective policy learning and leads to higher rewards.
[0058] In a possible implementation of the first aspect, the step of randomly sampling from the common experience pool and the LLM experience pool according to a preset sampling ratio includes:
[0059] In 0.5< Generate a random number in the interval <1 ;
[0060] Randomly drawn from the common experience pool Ordinary experience data, randomly drawn from the LLM experience pool empirical data based on LLM correction; among them, is the number of samples in a batch;
[0061] The sampled ordinary experience data and the sampled LLM-corrected experience data are randomly combined to form the sampled experience data.
[0062] In this approach, experience data corrected by the large language model is prioritized at a higher sampling rate, allowing the agent to focus on high-quality experience. Because the LLM experience pool outperforms the standard experience pool, prioritizing training with LLM-corrected experience data from the LLM experience pool accelerates convergence. At the same time, the standard experience pool is retained with a lower probability, reducing the probability of overfitting in the policy network and ensuring diversity in the experience data.
[0063] In a possible implementation manner of the first aspect, the target Q value is expressed as:
[0064] ,
[0065] in, represents the immediate reward at time t, represents the Q value of the state-action pair at time t+1 predicted by the value network, represents the entropy coefficient, represents the entropy term.
[0066] In a possible implementation of the first aspect, the step of training the agent by minimizing a loss function based on the normal experience pool and the LLM experience pool further includes:
[0067] Obtaining a second control action from the sampled LLM-corrected empirical data, and inputting the sampled LLM-corrected empirical data into the policy network to obtain a third control action predicted by the policy network;
[0068] The imitation learning loss is calculated according to the second control action, the third control action and the imitation learning loss function. The calculation expression of the imitation learning loss is:
[0069] ,
[0070] in, represents the imitation learning loss, represents empirical data based on LLM correction, represents the third control action obtained according to the strategy generated by the strategy network, represents the second control action from the empirical data based on the LLM correction;
[0071] The total loss of the policy network is calculated based on the imitation learning loss and the standard policy network loss. The expression of the total loss of the policy network is:
[0072] ,
[0073] in, represents the total loss of the policy network, represents the weight coefficient;
[0074] The network parameters of the policy network are updated by minimizing the total loss of the policy network.
[0075] In the above scheme, imitation learning is introduced during the training process of the intelligent agent to align the agent's actions with the actions in high-quality experience data, so that the intelligent agent can explore more valuable areas in the state space and reduce ineffective exploration, thereby further improving the efficiency and accuracy of deep reinforcement learning in building energy control.
[0076] In a second aspect, an embodiment of the present application provides a deep reinforcement learning system for building energy control, comprising:
[0077] A first building module is configured to build an intelligent agent based on a target building environment; the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control actions performed on the target building environment, and the reward function represents the reward value obtained after performing the control actions under the environmental state;
[0078] a collection module, configured to collect common experience data each time the intelligent agent interacts with the target building environment, and store the common experience data in a common experience pool;
[0079] A second construction module is configured to construct a prompt text based on the state space, the action space, and the intention description text; the intention description text is configured to indicate an action range of the control action that can be executed under each of the environmental states;
[0080] An acquisition module, configured to input the prompt text into a large language model to obtain an action range of each control action under each environmental state;
[0081] a correction module, configured to correct the common experience data according to the motion range to obtain experience data corrected based on the LLM, and store the experience data corrected based on the LLM in an LLM experience pool;
[0082] A training module is used to train the intelligent agent by minimizing the loss function according to the common experience pool and the LLM experience pool to obtain a trained building energy control strategy model.
[0083] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the deep reinforcement learning method for building energy control described in any one of the first aspects above is implemented.
[0084] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the deep reinforcement learning method for building energy control described in any one of the first aspects above.
[0085] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a terminal device, enables the terminal device to execute the deep reinforcement learning method for building energy control described in any one of the above-mentioned first aspects.
[0086] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0088] Figure 1 This is a flowchart of a deep reinforcement learning method for building energy control provided by an embodiment of the present application;
[0089] Figure 2 This is a schematic diagram of the composition of the prompt text provided in one embodiment of the present application;
[0090] Figure 3 This is a schematic diagram of a framework of a deep reinforcement learning method for building energy control provided by an embodiment of the present application;
[0091] Figure 4 This is a comparison chart of indoor temperature curves of the present method embodiment and the traditional method provided in one embodiment of the present application;
[0092] Figure 5 This is a comparison diagram of the reward curve convergence process between the present method embodiment provided in one embodiment of the present application and the traditional method;
[0093] Figure 6 Schematic diagram of the structure of a deep reinforcement learning system for building energy control provided by an embodiment of the present application;
[0094] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0095] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0096] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0097] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0098] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0099] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0100] Figure 1 The following is a schematic flow chart of a deep reinforcement learning method for building energy control provided by the present application. As an example and not a limitation, the method includes the following steps:
[0101] S11. Build an intelligent agent based on the target building environment.
[0102] Among them, the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control action performed on the building energy system of the target building environment, and the reward function represents the reward value obtained after performing the control action under the environmental state.
[0103] S12. Collect common experience data each time the agent interacts with the target building environment, and store the common experience data in a common experience pool.
[0104] S13. Construct prompt text based on the state space, action space and intention description text.
[0105] The intent description text is used to indicate the action range of control actions that can be executed in each environmental state.
[0106] S14. Input the prompt text into the large language model to obtain the action range of each control action in each environmental state.
[0107] S15 , correcting the common experience data according to the range of motion to obtain experience data corrected based on the LLM, and storing the experience data corrected based on the LLM in the LLM experience pool.
[0108] S16. Based on the normal experience pool and the LLM experience pool, the intelligent agent is trained by minimizing the loss function to obtain a trained building energy control strategy model.
[0109] The deep reinforcement learning method for building energy control provided in this embodiment leverages the prior knowledge embedded in a large language model to analyze various environmental states and infer an adaptive action range based on control intent. This guides and corrects low-quality experience data during the agent's random exploration, reducing ineffective exploration. By constructing a dual experience pool that stores both random and high-quality experience data, combined with a deep reinforcement learning algorithm, the agent can quickly and effectively explore the optimal control strategy, improving the efficiency and overall performance of building energy control strategy learning.
[0110] In one possible implementation, the target building environment in this embodiment represents a building that includes a DHW (Domestic Hot Water) system, a battery energy storage system, and an HVAC (Heating, Ventilation, and Air Conditioning) system. The agent's control task involves analyzing environmental conditions to adjust the HVAC system's heat pump power and the charge and discharge ratios of the DHW and battery energy storage systems.
[0111] In one possible implementation, the agent can be trained in the CityLearn simulation environment. CityLearn is an open-source Gym platform (reinforcement learning simulation environment) that provides a simplified building energy model for simulating energy management in a grid-interactive building community. The building's power comes from the grid, photovoltaic panels, and batteries. The building dataset in CityLearn can simulate a variety of building energy control scenarios with different weather conditions, building characteristics, and energy system settings.
[0112] Optionally, an implementation of step S11 specifically includes:
[0113] S111. Construct a state space based on the detection data of the target building environment.
[0114] In one possible implementation, the state space represents a continuous building space containing observations from four categories: calendar-level states, weather-level states, neighborhood-level states, and building-level states, such as the current time of day, outdoor dry-bulb temperature, carbon intensity, and cooling demand. The state space is expressed as:
[0115] ,
[0116] in, represents the environmental state at time t, Indicates the calendar status. Indicates weather conditions. represents the community-level status, Indicates building-level status.
[0117] S112. Construct an action space based on building energy system data of the target building environment.
[0118] In one possible implementation, the action space represents a continuous action space that determines the changes in building energy system data such as the charge and discharge ratios of the DHW (hot water storage system) and battery energy storage, and the output power of the heat pump. The expression of the action space is:
[0119] ,
[0120] in, represents the control action executed at time t, represents the charge and discharge ratio of the hot water energy storage system, Expressed as the charge and discharge ratio of battery energy storage, Expressed as the output power of the heat pump.
[0121] S113. Taking any environment state in the state space as the current environment state, and determining a state transition probability based on any control action in the action space.
[0122] Specifically, the state transition probability represents the probability of transitioning to the next state after taking a certain action in the current state. The state transition probability defines how the target building environment updates its state based on the actions taken by the agent, that is, predicting the future state of the environment. The expression of the state transition probability is:
[0123] ,
[0124] in, represents the state transition probability, represents the environmental state at time t+1, represents the environmental state at time t, Represents the control action executed at time t.
[0125] S114. Construct a reward function based on the immediate reward and discount factor for executing any control action in the action space under the current environment state.
[0126] In one possible implementation, the goal of building energy control is to learn an effective control strategy (based on the state Dynamically adjust its control actions ), to maximize the cumulative reward. The reward function is expressed as:
[0127] ,
[0128] in, Indicates cumulative rewards, represents the number of training steps, represents the discount factor that balances current rewards and future rewards at time t, Indicates immediate reward, represents the environmental state at time t, Represents the control action executed at time t.
[0129] As an example and not a limitation, the total number of training steps for the agent in this embodiment is 40,000, and the discount factor The configuration is 0.99.
[0130] S115. Construct an intelligent agent based on the state space, action space, state transition probability and reward function.
[0131] The deep reinforcement learning method for building energy control provided in this embodiment models the building energy control problem as a Markov decision process, and designs an intelligent agent based on the state space, action space, and reward function in the Markov decision process. This method does not require any prior information about uncertain parameters or a clear building thermal dynamics model, and has wider applicability.
[0132] As an example and not a limitation, in one implementation of step S111, the state space is as shown in Table 1:
[0133] Table 1 State space
[0134]
[0135] Optionally, in one implementation of step S114, the instant reward is based on indicators such as indoor temperature deviation, building energy consumption, and grid ramping value (power ramp rate, the rate of change of power load over time). Lower deviation, less energy consumption, and smoother ramping fluctuations will result in higher rewards, thereby promoting comfort, efficiency, and grid reliability. Specifically, constructing the instant reward may include the following steps:
[0136] S1141. Construct a comfort bonus based on the indoor temperature of a target building environment.
[0137] In one possible implementation, the comfort bonus It reflects the deviation between the indoor temperature and the set comfort temperature, representing the comfort level of the occupants. The expression of comfort reward is:
[0138] ,
[0139] in, Indicates comfort bonus, represents the indoor temperature at time t, Indicates the set comfort temperature. Indicates the set temperature deviation threshold. What is measured is the absolute value of the temperature deviation between the two. Reflects the acceptable comfort zone range. As an example and not a limitation, Set to 1℃, for example, if the standard comfort temperature is set to 23℃, a deviation threshold of 1℃ is acceptable, that is, the comfort zone range of the indoor temperature is 22℃ to 24℃. If the value is within the comfort zone, the reward is 0. Otherwise, exceeding the deviation threshold will result in a negative reward, i.e., a penalty value.
[0140] S1142. Construct a grid reward based on the grid volatility and grid power consumption of the target building environment.
[0141] In one possible implementation, to improve grid reliability and reduce energy consumption, the ramping value and grid power consumption value, which measure the fluctuation of grid power consumption, are used as grid rewards. It should be noted that the grid power consumption in this embodiment refers to the net energy consumption of the building's own power generation (such as solar power generation) minus the power imported from the grid. Specifically, the expression for the grid reward is:
[0142] ,
[0143] in, represents the grid reward, Indicates the grid volatility, The numerical range of ), represents the symbolic function, represents the power consumption of the grid at time t, represents the power consumption of the grid at time t-1, and Represents the weight coefficient.
[0144] In the grid reward, Indicates the ramping value of the power consumption fluctuation of the power grid. The ramping index is , It reflects that the power consumption of the grid in the next time step cannot be greater than that in the previous time step. When the power consumption of the grid suddenly increases, a penalty is imposed. Quantifying grid power consumption, grid power consumption , Used to indicate the sign of the parameter. When it is negative, it means that the building's own power generation has surplus power while meeting its own power consumption needs. For positive rewards; When it is positive, it means that the building's own power generation is insufficient and it needs to import power from the grid. It is a negative reward, which means that positive grid power consumption is penalized.
[0145] S1143. Construct instant rewards based on comfort rewards and grid rewards.
[0146] Specifically, The time step is Instant rewards when comfortable and grid rewards The expression of immediate reward is:
[0147] ,
[0148] in, and represents the weight coefficient, and Decide and the relative importance of .
[0149] The deep reinforcement learning method for building energy control provided in this embodiment uses multi-objective weighted rewards to decompose complex task objectives into multiple quantifiable objectives of comfort and energy consumption, and guides the intelligent agent to find a balance between multiple objectives through weight adjustment, avoiding a single objective dominating the strategy and ensuring the optimal overall performance of the strategy.
[0150] Optionally, an implementation of step S12 may include:
[0151] S121. The intelligent agent interacts with the target building environment at the current moment to obtain the current environment state, and uses the intelligent agent to select and execute a first control action to obtain the next environment state returned by the target building environment and the reward value obtained by the intelligent agent for executing the first control action.
[0152] S122: Construct common experience data based on the current environment state, the first control action, the next environment state, and the reward value.
[0153] In one possible implementation, Indicates the interactions, the current environment state is represented by , the first control action is expressed as , the reward value is expressed as , the next environmental state is represented by , each piece of common experience data is represented by Composition, store each common experience data into the common experience pool.
[0154] Optionally, one implementation of step S13 may include: constructing a prompt text describing the background, goal, state space, action space, and output format of the building energy control so that the LLM (Large Language Model) can effectively understand the decision-making task. Specifically, the prompt text is as follows: Figure 2 As shown, Figure 2 STATE and ACTION can be replaced by the real-time environment state during the interaction (i.e. the current environment state ) and the required control action (i.e. the first control action The prompt text enables the LLM to decompose complex tasks into clearer sequential steps: first, the prompt LLM analyzes the environment state; second, the LLM is required to infer a reasonable range of actions based on the current environment state and the goal; finally, the LLM is required to output the range of control actions in the required output format.
[0155] Optionally, an implementation of step S14 may include: after constructing the prompt text, inputting it into an LLM model , from which all actions are generated, specifically:
[0156] ,
[0157] in, It is Interactive prompts, It is the reference range of all actions, consisting of multiple lists, each list corresponds to the reference range of a specific action, represents the total number of actions, and Specify the The lower and upper bounds of the range of motion.
[0158] As an example, not a limitation, GPT-4o-mini (a lightweight, optimized version of the Large Language Model) is used as the Large Language Model (LLM). Its powerful reasoning capabilities enable it to reason about the reasonable range of actions in complex building energy environments. Of course, this embodiment can also use GPT-4 (the fourth-generation general large language model) and GPT-4o (an optimized version of the Large Language Model) for different application scenarios.
[0159] Optionally, an implementation of step S15 may include:
[0160] S151: Obtain a target action range of a first control action.
[0161] Specifically, for the The first control action of the interaction , and its corresponding target action range is ,
[0162] S152: Correct the first control action to a second control action that meets the target action range.
[0163] In one possible implementation, the jth specific action be constrained within their corresponding reasonable reference ranges, specifically: ,
[0164] in, The function will Constrain to reference range middle, is the corrected action. The corrected actions constitute the LLM-guided action vector (ie the second control action).
[0165] ,
[0166] S153: Construct LLM-corrected experience data based on the current environment state, the second control action, the next environment state, and the reward value.
[0167] In one possible implementation, the empirical data corrected by the large language model is Composition, each piece of experience data based on LLM correction is stored in the LLM experience pool.
[0168] As an example and not by way of limitation, the LLM experience pool is constructed only in the time steps of the previous stage, that is, in approximately two thousand interactions. In this embodiment, the ordinary experience pool and the second experience recycling together constitute an experience recycling pool. As an example and not by way of limitation, an experience recycling pool of 10,000 is used to store experience data. Among them, the experience recycling pool enables the intelligent agent in deep reinforcement learning training to achieve rapid learning convergence based on a large amount of historical experience data. In one possible implementation, during the deep reinforcement learning training process, the experience recycling pool can be expanded and iterated, so that the intelligent agent can obtain comprehensive and valuable experience data.
[0169] In one possible implementation, the intelligent agent of this embodiment uses the Actor Network-Critic Network (Actor Critic) framework, which combines the policy gradient method and the value function method. That is, the intelligent agent is trained and learned through a policy network (Actor Network) for selecting actions and another value network (Critic Network) for evaluating the value of actions.
[0170] As an example and not a limitation, the Actor network and Critic network of the reinforcement learning agent are designed as two fully connected layers, each containing 256 neurons.
[0171] Optionally, an implementation of step S16 may include:
[0172] S161. Randomly sample from the common experience pool and the LLM experience pool according to a preset sampling ratio to obtain a batch of sampling experience data.
[0173] S162. Calculate the standard policy network loss based on the sampled empirical data and the policy network loss function.
[0174] In one possible implementation, the standard policy network loss is expressed as;
[0175] ,
[0176] in, represents the standard strategy network loss, represents the network parameters of the policy network, represents the computational expectation, Indicates from Each piece of sampled experience data, Indicates the experience data of the ordinary experience pool and the LLM experience pool. represents the entropy coefficient of the policy network, represents the policy generated by the policy network, represents the predicted Q value of the value network, represents the network parameters of the value network, represents the environmental state at time t, Represents the control action executed at time t.
[0177] S163. Update the network parameters of the policy network by minimizing the standard policy network loss.
[0178] In one possible implementation, stochastic gradient descent is used to minimize the standard policy network loss To update the policy network, specifically:
[0179] ,
[0180] in, represents the learning rate, yes About the network parameters of the policy network The gradient of the updated policy network is As an example and not a limitation, the learning rate in this embodiment is set to 0.0001.
[0181] S164. Calculate the value network loss based on the sampled empirical data and the value network loss function.
[0182] In one possible implementation, the value network loss is expressed as:
[0183] ,
[0184] in, represents the value network loss, represents the network parameters of the value network, represents the target Q value, represents the computational expectation, Indicates from Each piece of sampled experience data, Indicates the experience data of the ordinary experience pool and the LLM experience pool. represents the predicted Q value of the value network, represents the environmental state at time t, Represents the control action executed at time t.
[0185] In one possible implementation, the target Q value is expressed as:
[0186] ,
[0187] in, represents the target Q value, Indicates immediate reward, According to the policy network The action at time t+1 obtained by the generated strategy , is estimated by the value network State-action pair at a moment Q value, represents the entropy coefficient, Represents the entropy term, which is calculated by the policy network.
[0188] S165. Update the network parameters of the value network by minimizing the value network loss.
[0189] In one possible implementation, stochastic gradient descent is used to minimize the value network loss To update the value network, specifically:
[0190] ,
[0191] in, represents the learning rate, yes About the network parameters of the value network The gradient of the updated value network parameter is .
[0192] This embodiment provides a deep reinforcement learning method for building energy control. It achieves efficient learning through two collaborative networks: a standard policy network loss guides the policy network to optimize action selection, while a value network loss guides the value network to accurately estimate action values. These two methods work together to reduce gradient variance and improve learning stability. Furthermore, experience data is sampled from two separate experience pools and randomly shuffled, ensuring both data diversity and high-value experience for the reinforcement learning agent. This promotes effective policy learning and leads to higher rewards.
[0193] Optionally, an implementation of step S161 specifically includes:
[0194] S1611, in 0.5< Generate a random number in the interval <1 .
[0195] During training, priority sampling is applied. As an example and not a limitation, the sampling ratio If set to 0.6, is 0.4.
[0196] S1612, randomly drawn from the general experience pool Ordinary experience data, randomly drawn from the LLM experience pool Empirical data based on LLM correction.
[0197] in, is the number of samples in a batch. As an example but not a limitation, the size of N is set to 256, that is, each training uses 256 time steps of experience data. In one possible implementation, the common experience data sampled from the common experience pool is recorded as ,in , express The amount of empirical data collected in the formula At the same time, the LLM-corrected experience data sampled from the LLM experience pool is recorded as ,in , express The amount of empirical data collected in the formula Calculated.
[0198] S1613. Randomly combine the sampled common experience data and the sampled experience data based on LLM correction to obtain a batch of sampled experience data.
[0199] Specifically, and Connect them to form the sampled experience data sampled from the experience recycling pool:
[0200] ,
[0201] In one possible implementation, The sampled experience data in will be randomly disrupted and then used to update the network parameters of the policy network and the value network.
[0202] This embodiment provides a deep reinforcement learning method for building energy control. Experience data corrected by a large language model is prioritized at a higher sampling rate, enabling the agent to prioritize high-quality experience. Because the LLM experience pool outperforms the standard experience pool, prioritizing training with LLM-corrected experience data from the LLM experience pool accelerates convergence. At the same time, the standard experience pool is retained with a lower probability, reducing the probability of overfitting in the policy network and ensuring the diversity of the experience data.
[0203] Optionally, an implementation of step S16 further includes:
[0204] S166 , obtaining a second control action from the sampled LLM-corrected empirical data, and inputting the sampled LLM-corrected empirical data into the policy network to obtain a third control action predicted by the policy network.
[0205] In one possible implementation, define It is a strategic network LLM Experience Pool Middle The action predicted in the secondary environment state (the action predicted by the policy network is the third control action) is used represents the set of predicted third control actions, specifically:
[0206] ,
[0207] S167. Calculate the imitation learning loss based on the second control action, the third control action, and the imitation learning loss function.
[0208] In one possible implementation, define For LLM experience pool The first The LLM guided action under the secondary environmental state (the control action based on LLM correction is the second control action) is used Represents a set of second control actions for guidance, specifically:
[0209] ,
[0210] In one possible implementation, the imitation learning loss is constructed as the predicted action output by the policy network and LLM guided action The mean squared error (MSE) between them is used as a supervision mechanism to align the agent’s actions with the actions guided by the LLM. The calculation expression of the imitation learning loss is:
[0211] ,
[0212] in, represents the imitation learning loss, represents empirical data based on LLM correction, represents the third control action obtained according to the strategy generated by the strategy network, represents the second control action from the empirical data based on the LLM correction.
[0213] S168. Calculate the total loss of the strategy network based on the imitation learning loss and the standard strategy network loss.
[0214] The standard policy network loss is calculated in step S162 above and will not be described in detail here. In one possible implementation, the total policy network loss is expressed as:
[0215] ,
[0216] in, represents the total loss of the policy network, In one possible implementation, the weight controlling the degree of LLM guidance is initially set to a relatively high positive value and gradually reduced to 0 during training, so that the reinforcement learning agent can maximize the cumulative reward from the guidance of the LLM.
[0217] As an example and not a limitation, the total number of training steps for the agent is 40,000, and the imitation learning loss weight is It is set to 7 in the first 5000 steps to ensure that the reinforcement learning agent pays more attention to the guidance provided by the LLM guided experience. In the remaining training steps, and are all set to 0 so that the agent focuses on maximizing the cumulative reward.
[0218] S169. Update the network parameters of the policy network by minimizing the total loss of the policy network.
[0219] In one possible implementation, the stochastic gradient descent method is used to minimize the total loss of the policy network. To update the policy network, the process is:
[0220] ,
[0221] in, yes About parameters The updated parameters are .
[0222] This embodiment provides a deep reinforcement learning method for building energy control, which introduces imitation learning during the training process of the intelligent agent, aligning the agent's actions with actions in high-quality empirical data, thereby enabling the intelligent agent to explore more valuable areas in the state space and reduce ineffective exploration, thereby further improving the efficiency and accuracy of deep reinforcement learning in building energy control.
[0223] The overall framework of this embodiment is as follows Figure 3 As shown below, refer to Figure 3 Let's use an example. The state space includes time, indoor temperature, photovoltaic power generation, and cooling demand. The action space includes controlling the hot water storage charge and discharge ratio of water heaters, controlling the battery charge and discharge ratio of household appliances, and controlling the air conditioning power of air conditioners. A reinforcement learning agent interacts with the environment, generating actions based on the current state of the environment according to the agent's current policy. This generated experience data is then stored directly in a common experience pool. Simultaneously, a large language model (LLM) interacts with the environment to generate an action reference range, and corrects the actions generated by the current policy. This generates LLM-guided experience data, which is then stored in the LLM experience pool. Prioritized sampling allows the reinforcement learning agent to prioritize experience data from the LLM experience pool. The total actor network loss is constructed by combining the imitation learning loss and the original actor network loss. The actor network is trained by minimizing this total loss, while the critic network is trained by minimizing the critic network loss.
[0224] By way of example and not limitation, buildings B1, B2, B3, B4, and B5 from the CityLearn framework were selected for training and evaluation of this method embodiment. Specifically, building B1 represents a high-load commercial building with significant electricity and heat demand; building B2 represents a residential building with demand patterns influenced by residents' daily activities; building B3 represents an industrial building with energy demand concentrated during specific production periods; building B4 represents a mixed-use building with a complex demand pattern combining commercial and residential functions; and building B5 represents a green building equipped with efficient energy systems and integrated renewable energy.
[0225] Optionally, during the training phase, a dataset from building B1 is used. The dataset covers a month and a total of 720 data points, i.e., each data point corresponds to one hour of information about weather, building load demand, and number of users. During the evaluation phase, the agent is deployed in buildings B1 to B5 for generalization experiments.
[0226] As an example and not a limitation, the LLMPI agent represents an agent trained based on an embodiment of the present method, and the DRL baseline agent represents an agent trained without LLM guidance experience. The LLMPI agent and the DRL baseline agent are simulated in the CityLearn environment, both are trained using data from building B1, and both are deployed on building B1. Figure 4 The comparison of indoor temperature curves of the embodiment of the present method and the traditional method is shown. Figure 4 It can be seen that the LLMPI agent can control the indoor temperature to remain within the temperature comfort zone for more time, showing more effective temperature regulation. In contrast, the DRL baseline agent shows obvious fluctuations and often deviates from the comfort zone.
[0227] By way of example and not limitation, Figure 5 The comparison of the reward curve convergence process of this method embodiment and the traditional method is shown. Figure 5 As shown in the figure, the LLMPI agent achieved significantly higher rewards, demonstrating that LLMPI can more effectively learn multi-objective control strategies to ensure user comfort, improve grid reliability, and reduce energy consumption. Furthermore, the LLMPI agent achieved the maximum baseline reward of the DRL baseline agent before the LLMPI agent, indicating that LLMPI improves the efficiency of policy learning.
[0228] In one possible implementation, LLMPI agents and DRL baseline agents are deployed in different buildings for experiments, and their performance in different scenarios is evaluated using evaluation metrics. As an example and not a limitation, this embodiment uses four key metrics to evaluate the agent's ability to control user comfort, grid reliability, and reduce energy consumption in building energy control scenarios: Unmet Hours, Grid Power Consumption Ramping, 1-Load Factor, and Daily Peak. The smaller the values of these four evaluation metrics, the stronger the agent's ability. Specifically:
[0229] The Unmet Hours metric evaluates the proportion of time that the indoor temperature deviates from the set temperature and exceeds the acceptable threshold of the set temperature:
[0230] ,
[0231] in, Indicates the number of time steps in an evaluation episode. The index measures the ability of the agent to keep the indoor temperature within the comfort range. The lower the index value, the better the thermal comfort control.
[0232] The Ramping index is used to evaluate the fluctuation of electricity consumption over time and reflects the stability and smoothness of the electricity energy demand curve. It is expressed as:
[0233] ,
[0234] in, and The time steps are and Energy consumption when Indicates the number of time steps in an evaluation episode. The Ramping metric captures fluctuations in energy demand. A higher Ramping value indicates greater fluctuations in grid power consumption.
[0235] The 1-Load FActor network metric quantifies the volatility of electrical energy consumption by comparing the daily average consumption with the peak demand within a given period and is defined as follows:
[0236] ,
[0237] in, is the number of hours in a day (24 hours), Indicates the number of days, Is the evaluation round The number of evaluation days included in Indicates that the evaluation has reached the dth day in the total round, t represents the time step (one step is one hour), and the 1-Load FActor network calculation formula is 1 minus the ratio of daily average consumption to peak consumption. The lower the value, the more efficient and balanced the energy consumption.
[0238] The Daily Peak metric measures the average maximum energy consumption within a day. The formula is as follows:
[0239] ,
[0240] The Daily Peak value reflects the maximum power load a building draws from the grid each day. Smaller values indicate less energy demand fluctuation and lower grid stress, providing a measure of grid stability and operational reliability.
[0241] As an example and not a limitation, the LLMPI agent and the DRL baseline agent are deployed on buildings B1 to B5 in the CityLearn framework. The comparison results of the LLMPI agent and the DRL agent on the four evaluation metrics of Unmet Hours, Ramping, 1-Load FActor Network, and Daily Peak are shown in Table 2:
[0242] Table 2 Comparison results between LLMPI agent and DRL agent on four evaluation indicators
[0243]
[0244] It can be seen that the LLMPI agent outperforms the DRL baseline agent in most buildings and most indicators, demonstrating strong generalization ability.
[0245] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0246] Corresponding to the deep reinforcement learning method for building energy control described in the above embodiment, Figure 6 A structural block diagram of a deep reinforcement learning system for building energy control provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0247] Reference Figure 6 , the deep reinforcement learning system for building energy control includes:
[0248] The first construction module 11 is used to construct an intelligent agent based on the target building environment; the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control actions performed on the target building environment, and the reward function represents the reward value obtained after performing the control action under the environmental state.
[0249] The collection module 12 is used to collect common experience data each time the intelligent agent interacts with the target building environment, and store the common experience data in a common experience pool.
[0250] The second construction module 13 is used to construct a prompt text according to the state space, the action space and the intention description text; the intention description text is used to indicate the action range of the control action that can be executed in each environmental state.
[0251] The acquisition module 14 is used to input the prompt text into the large language model to obtain the action range of each control action in each environmental state.
[0252] A correction module 15 is configured to correct the common experience data according to the range of motion to obtain experience data corrected based on the LLM, and store the experience data corrected based on the LLM in the LLM experience pool;
[0253] The training module 16 is used to train the intelligent agent by minimizing the loss function based on the common experience pool and the LLM experience pool to obtain a trained building energy control strategy model.
[0254] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0255] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0256] Figure 7 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 7 As shown, the electronic device 2 of the embodiment includes: at least one processor 20 ( Figure 7 Only one is shown in the figure) a processor, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20. When the processor 20 executes the computer program 22, the steps in each of the above-mentioned embodiments of the deep reinforcement learning method for building energy control are implemented.
[0257] The electronic device 2 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will appreciate that Figure 7This is merely an example of the electronic device 2 and does not constitute a limitation on the electronic device 2 . The electronic device 2 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 2 may also include input and output devices, network access devices, etc.
[0258] The processor 20 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0259] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as the hard drive or memory of the electronic device 2. In other embodiments, the memory 21 may also be an external storage device of the electronic device 2, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 2. Furthermore, the memory 21 may include both an internal storage unit of the electronic device 2 and an external storage device. The memory 21 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory 21 may also be used to temporarily store data that has been output or is about to be output.
[0260] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps in the above-mentioned embodiments of the deep reinforcement learning method for building energy control.
[0261] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned embodiments of the deep reinforcement learning method for building energy control when executing the computer program product.
[0262] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0263] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0264] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A deep reinforcement learning method for building energy control, characterized in that include: An intelligent agent is constructed based on a target building environment; the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control action performed on the target building environment, and the reward function represents the reward value obtained after performing the control action under the environmental state; collecting common experience data each time the intelligent agent interacts with the target building environment, and storing the common experience data in a common experience pool; Constructing a prompt text according to the state space, the action space and the intention description text; the intention description text is used to indicate the action range of the control action that can be executed under each of the environmental states; Inputting the prompt text into a large language model to obtain an action range of each of the control actions in each of the environmental states; Correcting the common experience data according to the motion range to obtain LLM-corrected experience data, and storing the LLM-corrected experience data in an LLM experience pool; According to the common experience pool and the LLM experience pool, the intelligent agent is trained by minimizing a loss function to obtain a trained building energy control strategy model; The step of collecting common experience data each time the intelligent agent interacts with the target building environment comprises: The agent interacts with the target building environment at a current moment to obtain a current environment state, and uses the agent to select and execute a first control action to obtain a next environment state returned by the target building environment and a reward value obtained by the agent for executing the first control action; constructing the common experience data based on the current environment state, the first control action, the next environment state, and the reward value; The step of correcting the common experience data according to the motion range to obtain experience data corrected based on the LLM comprises: Acquiring a target action range of the first control action; Correcting the first control action to a second control action that meets the target action range; The LLM-corrected experience data is constructed based on the current environment state, the second control action, the next environment state, and the reward value.
2. The deep reinforcement learning method for building energy control according to claim 1, characterized in that The step of constructing an intelligent agent based on the target building environment includes: The state space is constructed according to the detection data of the target building environment. The expression of the state space is: , in, represents the environmental state at time t, Indicates the calendar status. Indicates weather conditions. represents the community-level status, Indicates building-level status; The action space is constructed according to the building energy system data of the target building environment. The expression of the action space is: , in, represents the control action executed at time t, represents the charge and discharge ratio of the hot water energy storage system, Expressed as the charge and discharge ratio of battery energy storage, Expressed as the output power of the heat pump; Taking any environmental state in the state space as the current environmental state, the state transition probability is determined based on any control action in the action space. The expression of the state transition probability is: , in, represents the state transition probability, represents the environmental state at time t+1; The reward function is constructed based on the immediate reward and discount factor for executing any control action in the action space under the current environment state. The expression of the reward function is: , in, Indicates cumulative rewards, represents the number of training steps, represents the discount factor that balances current rewards and future rewards, Indicates immediate reward; The intelligent agent is constructed based on the state space, the action space, the state transition probability and the reward function.
3. The deep reinforcement learning method for building energy control according to claim 2, characterized in that The deep reinforcement learning method for building energy control further includes constructing the immediate reward, and the step of constructing the immediate reward includes: A comfort reward is constructed according to the indoor temperature of the target building environment. The expression of the comfort reward is: , in, Indicates comfort bonus, represents the indoor temperature at time t, Indicates the set comfort temperature. Indicates the set temperature deviation threshold; A grid reward is constructed according to the grid volatility and grid power consumption of the target building environment. The expression of the grid reward is: , in, represents the grid reward, Indicates the grid volatility, The numerical range of ), represents the symbolic function, represents the power consumption of the grid at time t, represents the power consumption of the grid at time t-1, and represents the weight coefficient; The instant reward is constructed according to the comfort reward and the grid reward. The expression of the instant reward is: , in, and Represents the weight coefficient.
4. The deep reinforcement learning method for building energy control according to claim 1, characterized in that The intelligent agent corresponds to a policy network and a value network, and the step of training the intelligent agent by minimizing the loss function based on the common experience pool and the LLM experience pool includes: Random sampling is performed from the common experience pool and the LLM experience pool according to a preset sampling ratio to obtain a batch of sampling experience data; The standard policy network loss is calculated based on the sampled experience data and the policy network loss function. The expression of the standard policy network loss is: , in, represents the standard strategy network loss, represents the network parameters of the policy network, represents the computational expectation, Indicates from Each piece of sampled experience data, Indicates the experience data of the ordinary experience pool and the LLM experience pool. represents the entropy coefficient of the policy network, represents the strategy generated by the policy network, represents the predicted Q value of the value network, represents the network parameters of the value network, represents the control action executed at time t, represents the environmental state at time t; Updating the network parameters of the policy network by minimizing the standard policy network loss; The value network loss is calculated based on the sampled empirical data and the value network loss function. The expression of the value network loss is: , in, represents the value network loss, represents the target Q value; The network parameters of the value network are updated by minimizing the value network loss.
5. The deep reinforcement learning method for building energy control according to claim 4, characterized in that The step of randomly sampling from the common experience pool and the LLM experience pool according to a preset sampling ratio comprises: In 0.5< Generate a random number in the interval <1 ; Randomly drawn from the common experience pool Ordinary experience data, randomly drawn from the LLM experience pool empirical data based on LLM correction; among them, is the number of samples in a batch; The sampled ordinary experience data and the sampled LLM-corrected experience data are randomly combined to form the sampled experience data.
6. The deep reinforcement learning method for building energy control according to claim 4, characterized in that The expression of the target Q value is: , in, represents the immediate reward at time t, represents the discount factor, According to the policy network The action at time t+1 obtained by the generated strategy , represents the Q value of the state-action pair at time t+1 predicted by the value network, represents the entropy coefficient, represents the entropy term.
7. The deep reinforcement learning method for building energy control according to claim 4, characterized in that The step of training the agent by minimizing a loss function based on the common experience pool and the LLM experience pool further comprises: Obtaining a second control action from the sampled LLM-corrected empirical data, and inputting the sampled LLM-corrected empirical data into the policy network to obtain a third control action predicted by the policy network; The imitation learning loss is calculated according to the second control action, the third control action and the imitation learning loss function. The calculation expression of the imitation learning loss is: , in, represents the imitation learning loss, represents empirical data based on LLM correction, represents the third control action obtained according to the strategy generated by the strategy network, represents the second control action from the empirical data based on the LLM correction; The total loss of the policy network is calculated based on the imitation learning loss and the standard policy network loss. The expression of the total loss of the policy network is: , in, represents the total loss of the policy network, represents the weight coefficient; The network parameters of the policy network are updated by minimizing the total loss of the policy network.
8. A deep reinforcement learning system for building energy control, characterized in that include: A first building module is configured to build an intelligent agent based on a target building environment; the state space of the intelligent agent represents the environmental state of the target building environment, the action space represents the control actions performed on the target building environment, and the reward function represents the reward value obtained after performing the control actions under the environmental state; a collection module, configured to collect common experience data each time the intelligent agent interacts with the target building environment, and store the common experience data in a common experience pool; A second construction module is configured to construct a prompt text based on the state space, the action space, and the intention description text; the intention description text is configured to indicate an action range of the control action that can be executed under each of the environmental states; An acquisition module, configured to input the prompt text into a large language model to obtain an action range of each control action under each environmental state; a correction module, configured to correct the common experience data according to the motion range to obtain experience data corrected based on the LLM, and store the experience data corrected based on the LLM in an LLM experience pool; A training module, configured to train the intelligent agent by minimizing a loss function based on the common experience pool and the LLM experience pool to obtain a trained building energy control strategy model; The collection module is specifically configured to obtain a current environment state by the agent interacting with the target building environment at a current moment, and to use the agent to select and execute a first control action to obtain a next environment state returned by the target building environment and a reward value obtained by the agent when executing the first control action; and to construct the common experience data based on the current environment state, the first control action, the next environment state, and the reward value; The correction module is specifically configured to obtain a target action range of the first control action; correct the first control action to a second control action that meets the target action range; The LLM-corrected experience data is constructed based on the current environment state, the second control action, the next environment state, and the reward value.
Citation Information
Patent Citations
Building energy management method and device based on large model
CN120258330A
Building energy system with energy data stimulation for pre-training predictive building models
US20190378020A1