Intelligent Home Energy Management Method and System Based on Personalized Federated Reinforcement Learning
By adopting personalized federal reinforcement learning methods in smart home energy management, the problem of not considering the heterogeneity of the home energy system in the existing technology is solved, and an efficient and stable personalized energy management strategy is achieved, reducing energy costs and ensuring thermal comfort.
Patent Information
- Application Number
- CN202510407482.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing smart home energy management methods do not fully consider the heterogeneity of different home energy systems, resulting in the inability to obtain personalized energy management strategies with superior performance, and the existing methods are prone to overfitting and have weak adaptability.
Using a method based on personalized federated reinforcement learning, we use the method of modeling the energy cost minimization problem of heterogeneous smart families to design the environmental state, action and reward functions of Markov's decision-making process, and use the deep reinforcement learning algorithm optimization strategy in the edge-end agent to perform federated learning with the help of cloud central servers, obtain the pre-trained global model and perform personalized fine-tuning, and obtain an energy management strategy suitable for individual heterogeneous family environments.
Through personalized federal reinforcement learning methods, various smart families can share knowledge, solve overfitting problems, enhance the stability of energy management strategies, and consider heterogeneity to realize energy management strategies that adapt to their own environment, reduce energy costs and ensure thermal comfort.
Smart Images

Figure CN119903767B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent home energy management, and specifically relates to an intelligent home energy management method and system based on personalized federated reinforcement learning. Background Art
[0002] As an innovative form of modern power supply, the core feature of the smart grid is the deep integration of advanced information and communication technologies (such as the Internet of Things, sensor networks) into all links of energy production, transmission, distribution, and consumption. Under this technical framework, the intelligent home can effectively manage energy economy by dynamically optimizing the operation strategies of energy storage devices and controllable loads. It is worth noting that the heating, ventilation, and air conditioning (HVAC) system, which accounts for nearly 40% of the total residential energy consumption, shows significant energy-saving regulation value while maintaining the comfort of the indoor thermal environment. As a typical controllable load, the optimization of the operation mode of such temperature control devices not only needs to meet the human thermal comfort requirements, but also needs to find the optimal solution of the operation cost in the dynamic energy price system, so as to realize the coordinated optimization of residential comfort and energy economy.
[0003] In recent years, researchers have conducted in-depth studies on the problem of home energy management, which can be divided into model-based methods and learning-based methods according to whether they rely on system models. Traditional model-based energy management methods mainly rely on establishing an accurate model of the home energy system. However, in practice, due to the existence of various random factors, it is almost impossible to establish an accurate system model. Therefore, it is necessary to develop intelligent methods to more effectively manage home energy. Learning-based solutions have the advantage of relaxing the requirements for explicit system models, so they are another energy management strategy with more application prospects than traditional model-based methods. In particular, the model-free deep reinforcement learning method combines the advantages of reinforcement learning and deep learning, does not require environmental model information, and has achieved great success in the optimization decision-making of the smart grid.
[0004] Although significant progress has been made in smart home energy management at home and abroad, there are still some deficiencies. Existing research has not fully considered the heterogeneity existing in different household energy systems. Heterogeneity widely exists in the real world. Due to the differences in the parameters of different household energy systems (such as the rated power of the HVAC system, the rated charge and discharge power of the energy storage system, building structure, building materials, etc.), the dynamic models of each household's energy system will also be different, which in turn determines that the optimal energy management strategies of each household will be different. Existing research lacks personalized energy management strategies designed for environmental heterogeneity. Many algorithms proposed in the research let each household independently train the energy management strategy, and there are two deficiencies in the smart home energy management algorithm based on reinforcement learning / deep reinforcement learning: (1) Most existing algorithms use the data of the household itself for training, which is prone to overfitting, resulting in weak performance of the trained energy management strategy; (2) The strategy trained in a specific household environment is difficult to transfer to a new heterogeneous household environment, with weak adaptability. Summary of the Invention
[0005] In view of the deficiencies existing in the prior art, the present invention provides a smart home energy management method and system based on personalized federated reinforcement learning, aiming to make up for the deficiency that the existing smart home energy management method does not consider the heterogeneity of the household energy system, and solve the technical problem that the existing method cannot obtain a personalized energy management strategy with excellent performance.
[0006] To achieve the above object, the present invention is implemented by the following technical solutions: The present invention provides a smart home energy management method based on personalized federated reinforcement learning, including:
[0007] Step 1: Model the energy cost minimization problems of multiple heterogeneous smart home environments and design the environmental states, actions and reward functions corresponding to the Markov decision process;
[0008] Step 2: The edge agents in
[0009] heterogeneous smart home environments optimize the strategy locally using the deep reinforcement learning algorithm and perform federated learning with the help of the cloud central server to obtain a pre-trained global model with stable training performance; Step 3: Perform post-training fine-tuning on the edge agents in each heterogeneous smart home environment to obtain types of personalized energy management strategies applicable to
[0010] heterogeneous smart home environments;
[0011] Step 4: Deploy the personalized energy management strategies obtained by fine-tuning in the actual environment for operation. Furthermore, the expression of the energy cost minimization problem of the
[0012] ,
[0013] (1),
[0014] (2),
[0015] (3),
[0016] (4),
[0017] (5),
[0018] (6),
[0019] (7),
[0020] (8),
[0021] (9),
[0022] Wherein: represents the mathematical expectation, represents the time slot, represents at the electricity purchase cost from the public power grid in the time slot, represents at the depreciation cost of the energy storage system in the time slot; represents the charge and discharge power of the energy storage system in the time slot, represents charging, represents discharging; represents the input power of the HVAC system in the time slot; represents the electricity trading volume between the time slot and the public power grid, represents purchasing electricity from the main power grid, represents selling electricity to the main power grid; represents the electricity purchase price from the public power grid in the time slot, represents the electricity selling price to the public power grid in the time slot; represents the depreciation cost coefficient of the energy storage system; represents the energy level of the energy storage system in the time slot, represents the minimum energy level of the energy storage system, represents the maximum energy level of the energy storage system, Represents the maximum charging power of the energy storage system, Represents the maximum discharging power of the energy storage system, Represents The energy level of the time-slot energy storage system, Represents the charging efficiency coefficient of the energy storage system, Represents the discharging efficiency coefficient of the energy storage system; Represents the maximum input power of the HVAC system; Represents The photovoltaic power generation of the time slot, Represents The non-shiftable load power of the time slot; Represents The indoor temperature of the time slot, Represents The outdoor temperature of the time slot, Represents The random thermal disturbance of the time slot, Represents The indoor temperature of the time slot, Represents the unknown building thermal dynamic model; Represents the lower bound of the indoor comfort temperature; Represents the upper bound of the indoor comfort temperature; heterogeneous parameter set .
[0023] Furthermore, the environmental state and actions of the Markov decision process are as follows:
[0024] (10),
[0025] (11),
[0026] (12),
[0027] Wherein: Represents the environmental state of the heterogeneous smart home in The time slot, Is The relative time slot of the time slot on the current day, , Represents the action of the heterogeneous smart home in The time slot; Represents The reward obtained in the time slot, Represents The temperature deviation of the time slot, Represents the conversion coefficient from temperature deviation to cost.
[0028] Furthermore, in the step 2, the An edge - side agent network structure contains an actor network and a critic network , both of which are multi - layer neural networks. Among them: The actor network is a mapping from the environmental state to actions, parameterized using . The number of neurons in the input layer is aligned with the dimension of the environmental state , and the number of neurons in the output layer is aligned with the dimension of the actions . The activation function used in the hidden layer is the rectified linear unit function, and the hyperbolic tangent function is used in the output layer for range compression. The critic network is used to calculate the action value, parameterized using . The input is the concatenation of the environmental state and the action . The output is the action value of executing the action in the state . The corresponding number of neurons in the input layer is the sum of the state dimension and the action dimension, and the number of neurons in the output layer is 1. The hidden layer also uses the rectified linear unit function. The edge - side agent also has two target networks: the actor target network and the critic target network . The structures of the actor target network and the critic target network are the same as the corresponding actor network and critic network, and the parameters are cloned from the original network regularly.
[0029] Furthermore, in step 2, the local training process of the edge - side agent of the th heterogeneous smart home environment is as follows:
[0030] (1) The edge - side agent executes the policy to interact with the environment to collect experience tuples , and stores them in the experience buffer;
[0031] (2) Sample a small batch of experiences from the experience buffer to update the parameters of the critic network and the actor network of the edge - side agent. The loss function expressions of the critic network and the actor network are as follows:
[0032] (13),
[0033] (14),
[0034] Among them: represents the loss of the critic network of the edge - side agent, represents the loss of the actor network of the edge - side agent, is the critic network of the edge - side agent, represents the critic target network of the edge - side agent, is the discount factor of the Markov decision process, represents the edge agent actor network, represents the edge agent actor target network;
[0035] (3) Soft-update the parameters of the edge agent actor target network and the critic target network. The expression is as follows:
[0036] (15),
[0037] (16),
[0038] where: , represents a merging factor;
[0039] (4) Repeat the above steps (1) - step (3) times.
[0040] Furthermore, during the federated learning process in step 2, the parameter aggregation expression of the edge agents is as follows:
[0041] (17),
[0042] where: represents the network layer index, represents the layer parameter index, represents the th edge agent's th parameter of the th layer during aggregation, represents the th parameter of the th layer of the global model; The way of distributing the global model parameters is as follows: Each edge agent model directly clones the network parameters of the global model as the initial weights for its new round of policy training. The expression is as follows: (18).
[0043] Furthermore, during the post-training fine-tuning process, each node stops uploading the edge agent parameters to the central server, stops obtaining the global model parameters from the central server, restricts the parameter distance between its own edge agent network model and the pre-trained global model, reduces the learning rate, and performs personalized fine-tuning in its own heterogeneous environment until the performance reaches stability and then stops. By adding a penalty term to the edge agent network loss function of each node to constrain the parameter distance between the model and the global model. The expression is as follows:
[0044] (19),
[0045] Wherein: represents the new loss function after adding the distance penalty term, represents the original loss function, represents a scaling factor, represents any distance calculation method (such as , distance), represents the network parameters of the th edge - side intelligent agent, represents the global model parameters.
[0046] Furthermore, the training process of the personalized energy management strategy can be divided into two parts: The first part, the training process of the pre - trained global model is as follows:
[0047] (1) Initialize the network parameters of the edge - side intelligent agents of heterogeneous smart homes and the experience buffer ;
[0048] (2) In each heterogeneous smart home, the edge - side intelligent agent interacts with the environment to collect experience tuples and stores them in the experience buffer ;
[0049] (3) Sample experience tuples from the experience buffer and independently train the edge - side intelligent agent using the deep deterministic policy gradient algorithm;
[0050] (4) The edge - side intelligent agent uploads the network parameters to the cloud - center server, and after aggregation, the global model parameters are obtained;
[0051] (5) The cloud - center server distributes the global model parameters to the edge - side intelligent agents of each edge node;
[0052] (6) Repeat the above steps (2) - step (5) until a pre - trained global model with stable training performance is obtained;
[0053] The second part, the steps of the post - training fine - tuning process are as follows:
[0054] (1) The edge - side intelligent agent loads the trained pre - trained global model parameters and stops uploading parameters to the central server;
[0055] (2) Soft - constrain the parameter distance between the local model and the pre - trained global model by adding a penalty term, that is, the calculation method of the edge - side intelligent agent loss function changes from to ;
[0056] (3) The edge - side agent reduces the learning rate to the original , continues to interact with the environment to collect experiences, and samples experience tuples to fine - tune the parameters of the edge - side agent using the deep deterministic policy gradient algorithm until the performance reaches stability;
[0057] (4) Repeat the above steps (1) - (3) for heterogeneous smart home environments, and obtain personalized energy management strategies that are respectively adapted to
[0058] To solve the above problems, the present invention further includes a smart home energy management system based on personalized federated reinforcement learning, which includes the following modules:
[0059] Edge - side information collection module: used to collect the current state information of heterogeneous smart home environments for use by the edge - side learning module and the edge - side inference module;
[0060] Edge - side inference module: used to perform inference based on the current state information and the policy neural network, and output the current actions of the energy storage system and the HVAC system;
[0061] Edge - side action execution module: used to control the energy storage system and the HVAC system and handle abnormal actions according to the current actions;
[0062] Edge - side learning module: used to calculate the reward information of heterogeneous smart home environments according to the current state, current actions, and the state of the next time slot, store experience tuples, and perform local energy management strategy training;
[0063] Cloud - based central server: used to aggregate the model parameters of edge nodes and output global model parameters;
[0064] Communication module: used for sending and receiving information between the edge - side and the cloud - based central server, regularly sending the local edge - side agent model parameters of all edge nodes to the cloud - based central server, and broadcasting the global model parameters of the cloud - based central server to all edge nodes.
[0065] Further, its operation process steps are as follows:
[0066] (1) The edge - side information collection module collects the current state information of heterogeneous smart home environments and delivers it to the edge - side inference module and the edge - side learning module;
[0067] (2) The edge - side inference module infers actions according to the current state information and delivers them to the edge - side action execution module and the edge - side learning module;
[0068] (3) Edge - side action execution module: In the heterogeneous smart home environment, operate the energy storage system and the HVAC system according to actions. When the environment evolves to the next time slot, the edge - side information collection module collects the state information of the heterogeneous smart home environment again and delivers it to the learning module;
[0069] (4) Edge - side learning module: According to the current state, current action, and the state of the next time slot, calculate the obtained reward according to the designed reward function, store the experience tuple in the experience buffer, and sample a batch of experience tuples in the experience buffer to update the local edge - side agent model parameters;
[0070] (5) Repeat steps (1) - (4) times;
[0071] (6) Communication module: Send the local edge - side agent model parameters of all edge nodes to the cloud - center server;
[0072] (7) Cloud - center server: Aggregate the edge - side agent model parameters from edge nodes and output the global model parameters;
[0073] (8) Communication module: Broadcast the global model parameters to the edge - side inference module and the edge - side learning module of all edge nodes;
[0074] (9) Repeat steps (1) - (8).
[0075] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Compared with the existing method of independently learning energy management strategies, the method of the present invention enables each smart home to share knowledge, solves the over - fitting problem caused by insufficient experience samples of a single home, and enhances the stability of the energy management strategy training process; (2) Compared with the existing method based on traditional federated learning, the method of the present invention considers the heterogeneity existing in the smart home energy system. Through the collaborative optimization of global model pre - training and personalized local fine - tuning, each home can learn an energy management strategy adapted to its own heterogeneity and naturally has scalability, effectively reducing the energy cost while fully ensuring the user's thermal comfort. Brief Description of the Drawings
[0076] Figure 1 is the flow chart of the method of the present invention;
[0077] Figure 2 is the personalized federated reinforcement learning architecture diagram proposed by the method of the present invention;
[0078] Figure 3 is the energy cost comparison diagram between the method of the present invention and other methods;
[0079] Figure 4 is the temperature deviation comparison diagram between the method of the present invention and other methods. Detailed implementation manners
[0080] To clearly illustrate the technical solution and innovative advantages of the present invention, the specific implementation manners of the present invention will be described below in conjunction with the accompanying drawings. It should be particularly noted that this implementation case is only for explanatory purposes and does not constitute a limitation on the claims of the present invention. This solution proposes a smart home energy management method based on personalized federated reinforcement learning. The energy system consists of a distributed photovoltaic power generation device, a battery energy storage system, a load unit, and a smart home energy management system. Specifically, the battery energy storage system optimizes the home energy cost by dynamically storing surplus photovoltaic power and implementing energy feedback during power shortages. The load unit includes two types: rigid loads and controllable loads. Rigid loads refer to necessary electrical appliances that maintain the basic living needs of users (such as televisions, microwave ovens, etc.), and controllable loads specifically refer to a heating, ventilation, and air conditioning (HVAC) system with power adjustment capabilities. The smart home energy management system performs hourly dynamic optimization decisions on the charge and discharge power of the energy storage system and the input power of the HVAC system based on the real-time operating parameters obtained (including but not limited to photovoltaic power generation, dynamic electricity price signals, ambient temperature, and rigid load demand), so as to achieve multi-energy collaborative control and maximize economic benefits. As Figure 1 shown, the smart home energy management method based on personalized federated reinforcement learning includes: Step 1: Model the energy cost minimization problems of multiple heterogeneous smart homes and design the environmental states, actions, and reward functions corresponding to the Markov decision process. Step 2: Edge agents in the heterogeneous smart home environments locally optimize the strategy using a deep reinforcement learning algorithm and perform federated learning with the help of a cloud central server to obtain a pre-trained global model with stable training performance. Step 3: Perform post-training fine-tuning on the edge agents in each heterogeneous smart home environment to obtain types of personalized energy management strategies applicable to heterogeneous smart home environments. Step 4: Deploy the personalized energy management strategies obtained through fine-tuning in the actual environment for operation.
[0081] Furthermore, the expression of the energy cost minimization problem for the th heterogeneous smart home environment is as follows:
[0082] ,
[0083] (1),
[0084] (2),
[0085] (3),
[0086] (4),
[0087] (5),
[0088] (6),
[0089] (7),
[0090] (8),
[0091] (9),
[0092] where: represents the mathematical expectation, represents the time slot, represents at the electricity purchase cost from the public power grid during the time slot, represents at the depreciation cost of the energy storage system during the time slot; represents the charge and discharge power of the energy storage system during the time slot, represents charging, represents discharging; represents the input power of the HVAC system during the time slot; represents the electricity transaction volume between the time slot and the public power grid, represents purchasing electricity from the main power grid, represents selling electricity to the main power grid; represents the electricity purchase price from the public power grid during the time slot, represents the electricity selling price to the public power grid during the time slot; represents the depreciation cost coefficient of the energy storage system; represents the energy level of the energy storage system during the time slot, represents the minimum energy level of the energy storage system, represents the maximum energy level of the energy storage system, represents the maximum charging power of the energy storage system, represents the maximum discharging power of the energy storage system, represents the energy level of the energy storage system during the time slot, represents the charging efficiency coefficient of the energy storage system, represents the discharging efficiency coefficient of the energy storage system; represents the maximum input power of the HVAC system; represents the photovoltaic power generation during the time slot, represents The non-time-shiftable load power of the time slot; denote The indoor temperature of the time slot, denote The outdoor temperature of the time slot, denote The random thermal disturbance of the time slot, denote The indoor temperature of the time slot, denote the unknown building thermal dynamic model; denote the lower bound of the indoor comfort temperature; denote the upper bound of the indoor comfort temperature; the heterogeneous parameter set .
[0093] Furthermore, the environmental state and actions of the Markov decision process are as follows:
[0094] (10),
[0095] (11),
[0096] (12),
[0097] Where: denote the environmental state in the heterogeneous smart home at the is the relative time slot of the time slot on the current day, , denote the action of the heterogeneous smart home at the time slot; represent the reward obtained at the denote the temperature deviation at the denote the conversion coefficient from the temperature deviation to the cost.
[0098] Furthermore, the th edge-side agent network structure in step 2 includes an actor network and a critic network , both of which are multi-layer neural networks, where: the actor network is a mapping from the environmental state to the action, parameterized using , the number of neurons in the input layer is aligned with the dimension of the environmental state, the number of neurons in the output layer is aligned with the dimension of the action, the activation function used in the hidden layer is the rectified linear unit function, and the hyperbolic tangent function is used in the output layer for range compression. The critic network is used to calculate the action value, using Parameterize with the input being the environmental state and the action concatenated. The output is the action value of executing the action in the state corresponding to the sum of the state dimension and the action dimension for the number of neurons in the input layer, 1 for the number of neurons in the output layer, and the rectified linear unit function is also used in the hidden layer. The edge agent also has two target networks: the actor target network and the critic target network . The actor target network and the critic target network have the same structure as the corresponding actor network and critic network, and the parameters are periodically cloned from the original network.
[0099] Furthermore, in step 2, the local training process of the edge agent of the th heterogeneous smart home environment is as follows:
[0100] (1) The edge agent executes the policy to interact with the environment to collect experience tuples and stores them in the experience buffer;
[0101] (2) Sample a small batch of experiences from the experience buffer to update the parameters of the critic network and actor network of the edge agent. The loss function expressions of the critic network and actor network are as follows:
[0102] (13),
[0103] (14),
[0104] where: represents the loss of the critic network of the edge agent, represents the loss of the actor network of the edge agent, is the critic network of the edge agent, represents the critic target network of the edge agent, is the discount factor of the Markov decision process, represents the actor network of the edge agent, represents the actor target network of the edge agent;
[0105] (3) Soft update the parameters of the actor target network and critic target network of the edge agent, and the expressions are as follows:
[0106] (15),
[0107] (16),
[0108] Wherein: , represents a merging factor;
[0109] (4) Repeat the above steps (1)-(3) times.
[0110] Furthermore, during the federated learning process of step 2, the parameter aggregation expression of the edge agents is as follows:
[0111] (17),
[0112] Wherein: represents the network layer index, represents the layer parameter index, represents the th parameter of the th layer of the th edge agent during aggregation, represents the th parameter of the th layer of the global model; the way of distributing the global model parameters is as follows: each edge agent model directly clones the network parameters of the global model as the initial weights for its new round of policy training, and the expression is as follows: (18).
[0113] Furthermore, during the post-training fine-tuning process, each node stops uploading the edge agent parameters to the central server, stops obtaining the global model parameters from the central server, restricts the parameter distance between its own edge agent network model and the pre-trained global model, reduces the learning rate, and performs personalized fine-tuning in its own heterogeneous environment until the performance reaches stability and then stops. By adding a penalty term to the network loss function of its own edge agent, the parameter distance between the model and the global model is constrained, and the expression is as follows:
[0114] (19),
[0115] Wherein: represents the new loss function after adding the distance penalty term, represents the original loss function, represents a scaling factor, represents any distance calculation method, represents the th network parameter of the edge agent, represents the global model parameter.
[0116] Furthermore, the training process of the personalized energy management strategy can be divided into two parts:
[0117] The training process of the pre-trained global model is as follows:
[0118] (1) Initialization The network parameters of the edge agents in heterogeneous smart homes and the experience buffer ;
[0119] (2) In each heterogeneous smart home, the edge agent interacts with the environment to collect experience tuples and stores them in the experience buffer ;
[0120] (3) Sample experience tuples from the experience buffer and independently train the edge agent using the Deep Deterministic Policy Gradient algorithm;
[0121] (4) The edge agent uploads the network parameters to the cloud center server, and after aggregation, the global model parameters are obtained;
[0122] (5) The cloud center server distributes the global model parameters to the edge agents of each edge node;
[0123] (6) Repeat the above steps (2)-(5) until a pre-trained global model with stable training performance is obtained;
[0124] The steps of the post-training fine-tuning process are as follows:
[0125] (1) The edge agent loads the trained pre-trained global model parameters and stops uploading the parameters to the central server;
[0126] (2) Soft-constrain the parameter distance between the local model and the pre-trained global model by adding a penalty term, that is, the calculation method of the edge agent loss function changes from to ;
[0127] (3) The edge agent reduces the learning rate to of the original, continues to interact with the environment to collect experience, and samples experience tuples to fine-tune the edge agent parameters using the Deep Deterministic Policy Gradient algorithm until the performance reaches stability;
[0128] (4) heterogeneous smart home environments repeat the above steps (1)-(3) to obtain types of personalized energy management strategies that are respectively adapted to heterogeneous environments.
[0129] To solve the above problems, the present invention includes a smart home energy management system based on personalized federated reinforcement learning, which comprises the following modules:
[0130] Edge - side information acquisition module: used to collect the current state information of the heterogeneous smart home environment for the edge - side learning module and the edge - side inference module;
[0131] Edge - side inference module: used to perform inference according to the current state information and the policy neural network, and output the current actions of the energy storage system and the HVAC system;
[0132] Edge - side action execution module: used to control the energy storage system and the HVAC system and handle abnormal actions according to the current actions;
[0133] Edge - side learning module: used to calculate the reward information of the heterogeneous smart home environment according to the current state, current actions and the state of the next time slot, store the experience tuples, and perform local energy management policy training;
[0134] Cloud - based central server: used to aggregate the model parameters of the edge nodes and output the global model parameters;
[0135] Communication module: used for sending and receiving information between the edge - side and the cloud - based central server, regularly sending the local edge - side agent model parameters of all edge nodes to the cloud - based central server, and broadcasting the global model parameters of the cloud - based central server to all edge nodes.
[0136] Furthermore, its operation process steps are as follows:
[0137] (1) The edge - side information acquisition module collects the current state information of the heterogeneous smart home environment and delivers it to the edge - side inference module and the edge - side learning module;
[0138] (2) The edge - side inference module infers actions according to the current state information and delivers them to the edge - side action execution module and the edge - side learning module;
[0139] (3) The edge - side action execution module operates the energy storage system and the HVAC system according to the actions in the heterogeneous smart home environment. The environment evolves to the next time slot, and the edge - side information acquisition module collects the state information of the heterogeneous smart home environment again and delivers it to the learning module;
[0140] (4) The edge - side learning module calculates the obtained reward according to the current state, current actions and the state of the next time slot according to the designed reward function, stores the experience tuple in the experience buffer, and samples a batch of experience tuples in the experience buffer to update the local edge - side agent model parameters;
[0141] (5) Repeat steps (1) - (4) times;
[0142] (6) The communication module sends the local edge agent model parameters of all edge nodes to the cloud center server;
[0143] (7) The cloud center server aggregates the edge agent model parameters from the edge nodes and outputs the global model parameters;
[0144] (8) The communication module broadcasts the global model parameters to the edge inference module and the edge learning module of all edge nodes;
[0145] (9) Repeat steps (1)-(8).
[0146] To demonstrate the performance of the method of the present invention, three comparative schemes are introduced.
[0147] Comparative scheme 1: A rule-based energy management strategy, and the adopted rules are split into two parts for setting. First, determine the power of the HVAC system at the current moment. When the indoor temperature exceeds the upper bound of the comfortable temperature, turn on the HVAC system at the maximum power; when the indoor temperature is lower than the lower bound of the comfortable temperature, turn off the HVAC system, otherwise keep the power of the previous time slot unchanged. Then determine the charge and discharge power of the energy storage system. When the power generation of the photovoltaic new energy system is greater than the power consumption of all loads in the smart home, store the excess power in the energy storage system. When the power generation of the photovoltaic new energy system is less than the power consumption of all loads (i.e., the HVAC system and conventional loads) in the smart home, the energy storage system releases electricity.
[0148] Comparative scheme 2: An energy management strategy based on deep reinforcement learning, that is, in each smart home, the deep deterministic policy gradient algorithm is used to independently learn the energy management strategy, and federated learning is not used for collaborative training.
[0149] Comparative scheme 3: An energy management strategy based on traditional federated deep reinforcement learning, that is, in each smart home, the deep deterministic policy gradient algorithm is used, and traditional federated learning is introduced to collaboratively train the energy management strategy. Instead of performing personalized fine-tuning, the energy management strategy corresponding to the global model is directly applied to all smart home environments.
[0150] Figure 3 and Figure 4 shows the average energy cost (the average value of the total energy cost of all households) and the average temperature deviation (the average value of the total temperature deviation of all households) of the method of the present invention and the comparative schemes in 20 smart home environments. As Figure 3 shown, compared with scheme 1 and scheme 2, the method of the present invention can reduce the average energy cost by 6.7%-15.7%. As Figure 4As shown, compared with all other solutions, the method of the present invention achieves the smallest average temperature deviation, effectively ensuring the thermal comfort of the household. In Comparative Solution 2, the energy management strategy is independently trained, and knowledge cannot be shared among smart homes. Both the average energy cost and the average temperature deviation indicators are inferior to the method of the present invention. Compared with Comparative Solution 3, the method of the present invention can reduce the average temperature deviation by 93.1% on the premise that the average energy cost increases by 0.2%.
[0151] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A smart home energy management method based on personalized federated reinforcement learning, characterized in that: The steps include: Step 1: Model the energy cost minimization problem of multiple heterogeneous smart homes and design the environment state, action and reward function corresponding to the Markov decision process; Step 2: The edge agents in N heterogeneous smart home environments use deep reinforcement learning algorithms to optimize strategies locally, and use the cloud central server to perform federated learning to obtain a pre-trained global model with stable training performance; Step 3: Post-train and fine-tune the edge agents of each heterogeneous smart home environment to obtain N personalized energy management strategies applicable to N heterogeneous smart home environments; In the post-training fine-tuning process of step 3, each node stops uploading edge agent parameters to the central server, stops obtaining global model parameters from the central server, limits the parameter distance between its own edge agent network model and the pre-trained global model, reduces the learning rate, and performs personalized fine-tuning in its own heterogeneous environment until the performance reaches stability. Each node adds a penalty term to its own edge agent network loss function to constrain the parameter distance between the model and the global model. The expression is as follows: in: Represents the new loss function after adding the distance penalty term, represents the original loss function, α represents a scaling factor, Represents any distance calculation method, ω n represents the network parameters of the nth edge agent, and ω represents the global model parameters; The learning process of personalized energy management strategy is as follows: In the first part, the training process of pre-trained global model is as follows: (1) Initialize the network parameters ω of the edge agents of N heterogeneous smart homes n And experience buffer (2) In each heterogeneous smart home, the edge agent interacts with the environment to collect experience tuples and stores them in the experience cache. (3) From the experience buffer The edge agent is trained independently using the deep deterministic policy gradient algorithm by sampling experience tuples in the dataset. (4) The edge agent uploads network parameters ω n To the cloud central server, after aggregation, the global model parameter ω is obtained; (5) The cloud center server sends the global model parameter ω to the edge agent of each edge node; (6) Repeat the above steps (2) to (5) until a pre-trained global model with stable training performance is obtained; In the second part, the post-training fine-tuning process steps are as follows: (1) The edge agent loads the pre-trained global model parameters ω and stops uploading the parameters to the central server; (2) By adding a penalty term, the parameter distance between the local model and the pre-trained global model is softly constrained, that is, the loss function calculation method of the edge agent is changed from Updated to (3) The edge agent reduces the learning rate to the original Continue to interact with the environment to collect experience, and sample experience tuples to fine-tune the parameters of the edge agent using a deep deterministic policy gradient algorithm until the performance reaches stability; (4) Repeat steps (1) to (3) for N heterogeneous smart home environments to obtain N personalized energy management strategies that are suitable for the N heterogeneous smart home environments; Step 4: Deploy the fine-tuned personalized energy management strategy in the actual environment.
2. The smart home energy management method based on personalized federated reinforcement learning according to claim 1 is characterized in that: In step 1, the expression of the energy cost minimization problem of the nth heterogeneous smart home environment is as follows: g n,t +p n,t =h n,t +e n,t +l n,t (8), in: represents mathematical expectation, t represents time slot, X n,1,t represents the cost of purchasing electricity from the public grid at time slot t, X n,2,t represents the depreciation cost of the energy storage system at time slot t; e n,t represents the charging and discharging power of the energy storage system in time slot t, e n,t >0 means charging, e n,t ≤0 means discharge; h n,t represents the input power of the HVAC system in time slot t; g n,t represents the amount of electricity traded with the public grid during time slot t, g n,t >0 means purchasing electricity from the main power grid, g n,t ≤0 means selling electricity to the main grid; represents the price of electricity purchased from the public grid at time slot t, represents the price of electricity sold to the public grid at time slot t; Represents the depreciation cost coefficient of the energy storage system; E n,t represents the energy level of the energy storage system in time slot t, represents the minimum energy level of the energy storage system, represents the maximum energy level of the energy storage system, Indicates the maximum charging power of the energy storage system, Indicates the maximum discharge power of the energy storage system, E n,t+1 represents the energy level of the energy storage system in the t+1 time slot, represents the charging efficiency coefficient of the energy storage system, Represents the discharge efficiency coefficient of the energy storage system; Indicates the maximum input power of the HVAC system; p n,t represents the photovoltaic power generation in time slot t, l n,t represents the non-time-shiftable load power in time slot t; represents the indoor temperature at time slot t, represents the outdoor temperature at time slot t, ξ n,t represents the random thermal disturbance in time slot t, represents the indoor temperature at time slot t+1, Represents the unknown thermal dynamic model of the building; Indicates the lower limit of indoor comfort temperature; represents the upper limit of indoor comfort temperature; heterogeneous parameter set 3. The smart home energy management method based on personalized federated reinforcement learning according to claim 1, characterized in that: In step 1, the environment state, action and reward function of the Markov decision process are designed as follows in the nth heterogeneous smart home: a n,t =(e n,t ,h n,t ) (13), r n,t =-(X n,1,t +X n,2,t +βX n,3,t ) (14), Where: s n,t represents the environmental state of the heterogeneous smart home at time slot t, t′ n is the relative time slot of time slot t on the day, t′ n =t%24;a n,t represents the action of the heterogeneous smart home in time slot t; r n,t represents the reward obtained in time slot t, X n,3,t represents the temperature deviation at time slot t, β represents the conversion coefficient from temperature deviation to cost, and p n,t represents the photovoltaic power generation in time slot t, l n,t represents the non-time-shiftable load power in time slot t, E n,t represents the energy level of the energy storage system in time slot t, represents the outdoor temperature at time slot t, Indicates the indoor temperature at time slot t.
4. The smart home energy management method based on personalized federated reinforcement learning according to claim 1 is characterized in that: The nth edge agent network structure in step 2 includes an actor network and a network of critics Both are multi-layer neural networks, where the actor network is a mapping from environment state to action, using θ n Parameterize the number of neurons in the input layer and the environment state s n,t Dimension alignment, the number of neurons in the output layer and action a n,t Dimensions are aligned, the activation function used in the hidden layer is the rectified linear unit function, the output layer uses the hyperbolic tangent function for range compression, and the critic network is used to calculate the action value using φ n Parameterized, the input is the environment state s n,t and action a n,t The output is in state s n,t Next, perform action a n,t The action value of the corresponding input layer neurons is the sum of the state dimension and the action dimension, the number of output layer neurons is 1, and the hidden layer also uses the rectified linear unit function. The edge agent also has two target networks: actor target network and critic target network Actor Target Network and critic target network The structure of is the same as the corresponding actor and critic networks, and the parameters are regularly cloned from the original networks.
5. The smart home energy management method based on personalized federated reinforcement learning according to claim 1 is characterized in that: In step 2, the local training process of the edge agent in the nth heterogeneous smart home environment is as follows: (1) Edge Agent Execution Strategy Interact with the environment to collect experience tuples (s n,t ,a n,t ,r n,t ,s n,t+1 ) and stored in the experience buffer, s n,t represents the environmental state of the heterogeneous smart home at time slot t, a n,t represents the action of the heterogeneous smart home in time slot t; r n,t Represents the reward obtained in time slot t; (2) Sample small batches of experience from the experience buffer and update the parameters of the critic network and actor network of the edge agent. The loss function expressions of the critic network and actor network are as follows: in: represents the loss of the edge agent critic network, represents the loss of the edge agent actor network, is the edge agent critic network, represents the edge agent critic target network, γ is the discount factor of the Markov decision process, represents the edge agent actor network, represents the target network of edge agent actors; (3) Soft update the parameters of the edge agent actor target network and the critic target network, expressed as follows: f′ n ←tf n +(1-τ)φ′ n (17), θ′ n ←tth n +(1-τ)θ′ n (18), Among them: τ<<1, τ represents a merging factor; (4) Repeat the above steps (1) to (3) Z times.
6. The smart home energy management method based on personalized federated reinforcement learning according to claim 1, characterized in that: In the process of federated learning in step 2, the parameter aggregation expression of N edge agents is as follows: Where: i represents the network layer index, j represents the layer parameter index, and w n,i,j represents the jth parameter of the i-th layer of the n-th edge agent during aggregation, ω i,j Represents the jth parameter of the i-th layer of the global model; the global model parameter delivery method is as follows: each edge agent model directly clones the network parameters of the global model as the initial weight of its own new round of strategy training, and the expression is as follows: oh n,i,j =ω i,j (20)。 7. The system used in the smart home energy management method based on personalized federated reinforcement learning according to any one of claims 1 to 6, characterized in that: The system includes the following modules: Edge information collection module: used to collect the current status information of the heterogeneous smart home environment for use by the edge learning module and edge reasoning module; Edge reasoning module: used to reason based on current state information and policy neural networks, and output the current actions of the energy storage system and HVAC system; Edge action execution module: used to control the energy storage system and HVAC system and handle abnormal actions according to the current action; Edge learning module: used to calculate the reward information of the heterogeneous smart home environment based on the current state, current action and next time slot state, store experience tuples, and perform local energy management strategy training; Cloud center server: used to aggregate the model parameters of edge nodes and output global model parameters; Communication module: used for sending and receiving information between the edge and the cloud center server, regularly sending the local edge agent model parameters of all edge nodes to the cloud center server, and broadcasting the global model parameters of the cloud center server to all edge nodes.
8. The system according to claim 7, characterized in that The operation process steps are as follows: (1) The edge information collection module collects the current status information of the heterogeneous smart home environment and delivers it to the edge reasoning module and the edge learning module; (2) The edge reasoning module infers the action based on the current state information and delivers it to the edge action execution module and the edge learning module; (3) The edge action execution module operates the energy storage system and HVAC system according to the action in the heterogeneous smart home environment. When the environment evolves to the next time slot, the edge information collection module collects the status information of the heterogeneous smart home environment again and delivers it to the learning module; (4) The edge learning module calculates the reward according to the designed reward function based on the current state, current action and next time slot state, stores the experience tuple in the experience buffer, and samples a batch of experience tuples in the experience buffer to update the local edge agent model parameters; (5) Repeat steps (1) to (4) J times; (6) The communication module sends the local edge agent model parameters of all edge nodes to the cloud central server; (7) The cloud center server aggregates the edge agent model parameters from the edge nodes and outputs the global model parameters; (8) The communication module broadcasts the global model parameters to the edge reasoning modules and edge learning modules of all edge nodes; (9) Repeat steps (1) to (8).
Citation Information
Patent Citations
Smart home energy management method and system based on deep reinforcement learning
CN110458443A