An intelligent home online energy management method combining prediction and learning
By combining model predictive control with deep reinforcement learning, an implicit world model of a smart home energy management system is constructed, which solves the problems of dependence on high-precision models and low sample utilization efficiency in existing technologies, and realizes efficient operation and cost optimization of the smart home energy management system.
Patent Information
- Application Number
- CN202511289737.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-10
AI Technical Summary
Existing smart home online energy management methods require explicit or high-precision indoor environmental thermodynamic models. Furthermore, deep reinforcement learning-based methods have low sample utilization efficiency, resulting in high training costs and deployment difficulties, making it hard to achieve long-term optimization and efficient environmental response.
By combining model predictive control with deep reinforcement learning, an implicit world model of a smart home energy management system is constructed. Data such as photovoltaic power generation, smart home rigid load, and real-time electricity price are obtained through white-box modeling. A Markov decision process is established to train the agent to make online decisions, reducing the reliance on high-precision models.
It achieves the goal of reducing the operating costs of smart homes while maintaining indoor thermal comfort, improving the long-term decision-making ability and sample utilization efficiency of intelligent agents, reducing the number of interactions with the real environment, and lowering training costs.
Smart Images

Figure CN120781716B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of smart home energy management and artificial intelligence, and specifically relates to a smart home online energy management method combining prediction and learning ideas. BACKGROUND
[0002] The continuous rise in global energy demand and the continuous development of renewable energy are driving the transformation of energy systems towards high efficiency, low carbon and intelligence. In this process, the development of smart home energy management systems has attracted widespread attention. Specifically, the smart home energy management system can improve the local consumption of renewable energy, reduce carbon emissions generated by the smart home purchasing electricity from the power grid, and reduce the energy cost of the smart home by integrating distributed renewable energy and energy storage systems and using advanced information communication and intelligent control technology while maintaining an indoor comfortable environment. Therefore, it is of great significance to study efficient and intelligent smart home online energy management methods.
[0003] At present, many smart home online energy management methods have been proposed to optimize the real-time operation of smart home energy management systems, mainly including model-based methods (such as model predictive control) and learning-based methods (such as deep reinforcement learning). Although the above model-based methods have made some progress, these methods require explicit or high-precision building indoor environment thermal dynamics models. Since the above models are related to many factors (such as outdoor environment, building material and real-time operation power of heating ventilation and air conditioning system), it is very challenging to build a high-precision building indoor environment thermal dynamics model. At the same time, even if a high-precision indoor environment thermal dynamics model is obtained, the existing model predictive control-based methods still have limitations, i.e., traditional model predictive control can only achieve short-term decision trajectory optimization and thus can only obtain a local optimal solution. Further, the existing deep reinforcement learning-based methods can avoid the above shortcomings, which do not require explicit indoor environment thermal dynamics models or predictions of uncertain parameters. However, the existing deep reinforcement learning-based methods have low sample utilization efficiency, requiring a large number of interactions with the real environment, greatly increasing the training cost and actual deployment difficulty.
[0004] The technical differences compared with the prior art are as follows:
[0005] Technical comparison with Chinese patent CN119903767A "Smart home energy management method and system based on personalized federated reinforcement learning";
[0006] I. Patent CN119903767A is aimed at multiple heterogeneous smart homes, using the technical architecture of "edge local deep reinforcement learning + cloud federated learning". First, the cloud center server carries out federated learning to obtain a pre-trained global model with stable training performance, and then fine-tunes the agents of each family edge, finally obtains N personalized energy management strategies. The core technology revolves around the combination of federated learning and deep reinforcement learning, and does not incorporate the prediction idea. This application focuses on a single smart home energy management system, the core of which is the combination of model predictive control (prediction idea) and deep reinforcement learning (learning idea). It does not rely on the federated learning architecture, not only builds an environment model for interaction with the agent, but also lets the agent actively learn the implicit world model related to the operation of the home energy management system, and focuses more on the collaborative optimization of prediction and learning.
[0007] II. Patent CN119903767A aims to "ensure occupant thermal comfort and reduce system energy cost" by adapting to multiple heterogeneous homes through the process of "global pre-training + personalized fine-tuning", but does not explicitly state the model support basis for online decision-making. This application first establishes a system operation cost minimization problem under the premise of maintaining indoor thermal comfort, then models it as a Markov decision process, and then builds an agent training framework based on the combination of model predictive control and deep reinforcement learning. The trained agent can make online decisions based on the learned implicit world model, and can simultaneously achieve the dual optimization of "reducing operating costs + improving indoor thermal comfort", with a more explicit model support for the decision-making process.
[0008] Comparison with Chinese patent CN110458443A "Smart Home Energy Management Method and System Based on Deep Reinforcement Learning";
[0009] I. Patent CN110458443A only uses single deep reinforcement learning technology, models the energy management problem as a Markov decision process without a building thermal dynamics model, designs the corresponding environment state, behavior and reward function, trains the optimal behavior of the energy storage system or controllable load through the deep deterministic policy gradient algorithm, and does not incorporate any prediction-related technology. This application breaks through the limitations of single learning technology, combines the dual advantages of model predictive control (prediction) and deep reinforcement learning (learning), and does not rely on the premise of "no building thermal dynamics model". Based on the modeling of the Markov decision process, an environment model for interaction with the agent is additionally constructed, forming a more complete technical system.
[0010] II. Patent CN110458443A adopts an online learning mode of "local testing combined with cloud training", and periodically copies the deep neural network parameters of cloud training to the local system to cope with environmental changes, and does not involve specific model assisted online decision mechanism. The present application establishes an intelligent agent training framework based on the constructed environment model, and the trained intelligent agent can directly act on the actual smart home system based on the obtained implicit world model for online decision, without the need for periodic parameter copying, and can further improve indoor thermal comfort while reducing operating costs through the cooperation of prediction and learning, and the response to environmental changes is more efficient and direct.
[0011] In summary, the existing smart home online energy management methods all have deficiencies, and it is urgent to research a smart home online energy management method that combines prediction and learning ideas, combines the dual advantages of model predictive control and deep reinforcement learning, and reduces the operating cost of the smart home while maintaining the indoor comfortable environment. SUMMARY
[0012] In view of the problems existing in the prior art, the present application proposes a smart home online energy management method that combines prediction and learning ideas, which aims to realize the joint optimization of the operating cost of the smart home energy management system and the indoor user thermal comfort, and to make up for the shortcomings of the existing model predictive control based method that needs to be clear or high precision indoor environment thermal dynamics model and can only realize short-term trajectory optimization, and to solve the technical problem that the existing deep reinforcement learning based method has low sample utilization efficiency and needs to interact with the real environment or accurate simulation environment a large number of times.
[0013] To achieve the above purpose, the technical scheme adopted by the present application is:
[0014] A smart home online energy management method that combines prediction and learning ideas, comprising the following steps:
[0015] (1) Under the premise of maintaining the indoor user thermal comfort, the problem of minimizing the operating cost of the smart home energy management system is established, including the objective function, the decision variable and the constraint condition;
[0016] (2) The problem of minimizing the operating cost is re-modeled as a Markov decision process related to the operation of the family energy management system, and an intelligent agent is constructed to learn the implicit world model related to the operation of the family energy management system;
[0017] (3) An environment model for the interaction of the intelligent agent learning the implicit world model of the family energy management system is constructed;
[0018] The environment model interacting with the agent learning the implicit world model of the smart home energy management system in step 3 comprises a battery energy storage system dynamic model and an indoor temperature dynamic model, and the above models are constructed by adopting a white box modeling method; data required for running the environment model comprises photovoltaic power generation power, smart home rigid load, outdoor temperature and real-time electricity price, and the above data are obtained from historical operation data of the smart home energy management system;
[0019] (4) According to the constructed environment model, an agent training framework based on combination of model predictive control and deep reinforcement learning is established to train the agent to obtain an implicit world model related to the operation of the home energy management system;
[0020] (5) After training, decision-making is performed based on the obtained implicit world model, and the decision-making is applied to the actual smart home system;
[0021] In step 5, decision-making is performed based on the obtained implicit world model, and the decision-making is applied to the actual smart home system, and the process only needs to call the encoder network and the policy network in the implicit world model of the home energy management system .
[0022] As a further improvement of the present application, the smart home energy management system operation cost minimization problem established in step 1 comprises an objective function, decision variables and constraint conditions, namely:
[0023]
[0024] Among them: is an expectation operator; and are energy cost and battery energy storage system depreciation cost, respectively; is the maximum time slot number of the home energy management system operation; is the home energy management system decision variable at the moment, wherein: and respectively represent the charging and discharging power of the battery energy storage system at the moment, is the input power of the heating, ventilation and air conditioning system at the moment; in formula (2), and are the charging efficiency and discharging efficiency of the battery energy storage system, respectively; in formula (3), , , are the minimum energy storage level, the maximum energy storage level and the energy storage level at the moment of the battery energy storage system, respectively; in formulas (4)-(5), and These are the maximum charging power and maximum discharging power of the battery energy storage system, respectively; Equation (6) indicates that charging and discharging of the battery energy storage system cannot occur simultaneously; in Equation (7), , , These are the inertia coefficient, performance coefficient, and overall thermal conductivity of the HVAC system. They represent The indoor and outdoor temperatures at any given time; in equation (8), and represent the upper and lower limits of the comfortable temperature range, respectively; in equation (9), The maximum input power of the HVAC system; in equation (10), , and They represent the first time. The electricity purchased from the grid at all times, the photovoltaic power generation capacity, and the electricity demand of rigid loads, including: Indicates that the smart home energy system is in the first At any given moment, electricity is sold to the grid; The smart home energy management system indicates that in the first... At any given time, electricity is purchased from the grid; in equation (11), and They represent the first time. The purchase and sale prices of electricity at any given time; in equation (12), This is the depreciation unit cost coefficient for battery energy storage systems.
[0025] As a further improvement of the present invention, the Markov decision process established in step 2 includes states related to the operation of the smart home energy management system. Actions related to the control of battery energy storage systems and HVAC systems And the reward function related to the operating costs of the smart home energy management system and indoor thermal comfort. The specific expression is as follows:
[0026]
[0027] in: Indicates the first The time is relative to the time within a 24-hour period of the day; Indicates the battery energy storage system in the first The charging and discharging power at any given time, if but . Weighting coefficients used to measure the importance of operating costs and deviations from indoor thermal comfort. This indicates the penalty for deviation from indoor thermal comfort, i.e. .
[0028] As a further improvement of the present application, the agent in step 2 comprises six neural networks, i.e. an encoder network , a dynamics network , a reward network , a policy network , a value function network set and a target value function network set, for constructing the latent world model related to the operation of the smart home energy management system. The above neural networks all adopt a deep neural network structure, which is composed of an input layer, multiple hidden layers and an output layer. Specifically, the encoder network , whose input is the state observation value of the smart home energy management system , outputs a low-dimensional latent state code The dynamics network , whose input is the current latent state , the final execution action , outputs the predicted next time latent state The reward network , whose input is the current latent state , the final execution action , outputs the predicted immediate reward The policy network , whose input is the current latent state , outputs the action mean value and logarithmic standard deviation conforming to the Gaussian distribution, which is used for sampling to generate the preliminary action The value function network set is composed of multiple independent multilayer perceptron integrations, and the input of each multilayer perceptron is the current latent state , the final execution action , and the output is the state-action value function value. The above neural networks constitute the latent world model related to the operation of the smart home energy management system, and the specific formula is as follows:
[0029] .
[0030] As a further improvement of the present application, the agent training framework based on the combination of model predictive control and deep reinforcement learning in step 4 specifically comprises the following steps:
[0031] First, initialize the neural network parameters of the agent, the experience replay pool and the interactive environment model, and then set the number of time slots for the current agent to interact with the environment model Finally, the preset iteration step is repeatedly executed until a termination condition is met.
[0032] As a further improvement of the application, the iteration step in step 4 comprises:
[0033] S51, at , it is determined whether the number of time slots in which the current intelligent agent interacts with the environment model reaches a preset maximum number of time slots for generating actions by random sampling , if the current number of time slots , random sampling is performed to obtain an action ; otherwise, the intelligent agent generates a final execution action based on a current environment observation value and an action planning function designed in combination with a model predictive control and deep reinforcement learning framework ;
[0034] S52, the intelligent agent executes the action in the constructed interactive environment to obtain a new state observation value and a reward function , then, an experience tuple , , , generated by the current interaction is stored in an experience replay pool ;
[0035] S53, it is determined whether the current interaction time slot reaches an update condition of each neural network model in the intelligent agent, if the update condition is reached, a small batch of data is sampled from the experience replay pool and the intelligent agent is updated based on a neural network update function combined with model predictive control and deep reinforcement learning;
[0036] S54, let , and it is determined whether the current interaction time reaches a preset maximum number of interaction time slots , if the current time , jump to step S51 until ; if the current time , terminate the training and obtain an implicit world model related to the operation of the smart home energy management system.
[0037] As a further improvement of the application, the action planning function designed in step S51 based on the model predictive control and deep reinforcement learning framework is specifically designed as follows:
[0038] S61, first, the encoder network in the smart home energy management system implicit world model reads the state observation value at time and map it to a low-dimensional latent state vector ;
[0039] S62, secondly, the smart home energy management system hidden world model calls a hybrid strategy to generate a set of candidate action sequences, each sequence has a preset length of , specifically, the strategy network of the smart home energy management system hidden world model first reads the latent state vector and performs step forward deduction to generate candidate action sequences, then, based on the Gaussian distribution random sampling generates candidate action sequences, wherein: and respectively represent the mean and standard deviation in Gaussian distribution sampling;
[0040] S63, then, based on the iterative model prediction path integral optimization algorithm, the candidate action sequences obtained in step S62 are continuously optimized;
[0041] S64, finally, after completing the optimization of the candidate action sequences, the output step of the final iteration of the action sequence is performed.
[0042] As a further improvement of the present application, the specific steps of continuously optimizing the candidate action sequences in step S63 based on the iterative model prediction path integral optimization algorithm are as follows:
[0043] S71, first, in each iteration, in addition to the action sequences generated by the strategy network in the home energy management system hidden world model, the remaining sequences obtained based on the random strategy are randomly resampled, for each candidate action sequence, starting from the current latent state , the dynamic network and the reward network in the hidden world model are used to execute the candidate action sequence to perform step forward simulation, and calculate the cumulative expected return of the simulation trajectory , the specific formula is as follows:
[0044]
[0045] wherein: is the latent state sequence, is the preset length of each sequence, is a discount factor; is the mean of the multiple independent multi-layer perceptrons in the value function network set;
[0046] S72, according to the cumulative expected return of the optimal action sequence obtained in step S71, the cumulative expected return of each candidate action sequence is calculated. From all candidate action sequences, the optimal sequence is selected to form an elite action sequence set.
[0047] S73, according to the elite action sequence set obtained in step S72, the weighted average value and the weighted standard deviation of are calculated, and the results are used to update the mean and the standard deviation of the Gaussian distribution in step S62.
[0048] S74, repeat the execution of the preset iteration number , and execute the action sequence output step of the final iteration.
[0049] As a further improvement of the present application, the output step of the final action sequence after the optimization of candidate action sequences in step S64 is as follows:
[0050] S81, according to the cumulative expected return of the final elite action sequence set obtained in step S64 , set the sampling weight of each elite action sequence, and randomly select an optimal action sequence from it using a differentiable discrete sampling method.
[0051] S82, according to the optimal action sequence obtained by sampling in step S81, the action at the first time step is extracted as the final output at the current decision moment. ;
[0052] S83, superimpose Gaussian noise sampled from on the action , wherein: is the standard deviation of the first time step .
[0053] As a further improvement of the present application, the specific steps of updating the agent in step S53 by sampling a small batch of data from the experience replay pool and updating the neural network function based on the combination of model predictive control and deep reinforcement learning are as follows:
[0054] S91, first, randomly select from the experience replay pool a starting sampling time index and extracts a time-continuous segment with a length of to obtain a small-batch agent update data;
[0055] S92, secondly, the reward value of the experience replay pool sampling and the observation value at the next moment are converted into the latent state at the next moment through an encoder network , and input into the policy network and the target value function network set to calculate the target value , and the specific calculation formula is as follows:
[0056]
[0057] wherein: is the output of the target value function network set, and is the minimum value of at least two independent multilayer perceptrons;
[0058] S93, then, the initial observation value is encoded into the initial latent state ; using the sampled action and the initial latent state , a multi-step evolution with a length of is performed in the dynamic space, a plurality of predicted latent states at the next moment are generated through a dynamic network , and the consistency loss between the predicted latent states at the next moment and the true encoded latent state is calculated , and the specific calculation formula is as follows:
[0059]
[0060] wherein: is a loss weighting factor, is a stop gradient operation function, and a predicted latent state sequence with a length of is obtained , denotes the predicted latent state at the next moment when t=0;
[0061] S94, according to the first latent state obtained in step S93 and the sampled action , the reward network and value network are used to predict the reward and TD error value The soft cross-entropy loss function is used to calculate the predicted reward. With real rewards The reward loss between, and Value and TD target value The value loss between them is calculated by weighting and summing all the above losses according to preset weighting coefficients to obtain the total loss function. The specific calculation formula is as follows:
[0062]
[0063] in: for Value function network parameters, Let cross-entropy be the loss function. For empirical tuples, The consistency loss coefficient, As the reward loss coefficient, This is the value loss coefficient;
[0064] S95. Based on the total loss function obtained in step S94, calculate the gradient using the backpropagation algorithm, and apply this gradient to the world model excluding the policy network. Optimize the parameters of all networks;
[0065] S96, The potential state sequence As input, by maximizing Value and policy entropy Calculate the policy loss and apply it to the policy network. parameters Independent optimization is performed, and the specific calculation formula is as follows:
[0066] (32)
[0067] in: and This is an adjustment factor for moving statistics, used for automatic balancing. Value and entropy term The relative size, for The mean of multiple independent multilayer perceptrons in a value function network;
[0068] S97. Update the target based on soft update method Value network parameters The specific formula is as follows:
[0069] (33)
[0070] in: This is the soft update coefficient.
[0071] Beneficial effects: A smart home online energy management method combining prediction and learning ideas, compared with the prior art, the present application combines the double advantages of model predictive control and deep reinforcement learning, and the beneficial effects achieved by the present application are as follows:
[0072] (1) Compared with the existing method based on model predictive control, the method of the present application does not need to know the explicit or high-precision indoor environment dynamic model, an implicit world model related to the operation of the smart home energy management system is constructed, the indoor thermal dynamics is accurately captured and the terminal value function is modeled, which is combined into the training process of the agent in the deep reinforcement learning algorithm, thereby improving the long-term decision-making ability of the agent and overcoming the limitation that the traditional model predictive control algorithm can only realize short-term trajectory optimization.
[0073] (2) Compared with the existing learning-based method, the method of the present application combines the model predictive control theory into the training process of the agent in the deep reinforcement learning algorithm, and through the rolling interaction with the implicit world model, the sample utilization efficiency of the agent training is significantly improved, without the need for a large number of interactions with the real environment or high-precision simulation environment, and has the advantage of low training cost. At the same time, after the training is completed, the method of the present application can obtain a control strategy with better performance and stability. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 is a flowchart of a smart home online energy management method combining prediction and learning ideas provided by the present application;
[0075] Figure 2 is a reward function training curve diagram of the method of the present application and other comparative schemes;
[0076] Figure 3 is a comparison diagram of the running cost of the home energy management system of the method of the present application and other comparative schemes;
[0077] Figure 4 is a comparison diagram of the average deviation of indoor comfortable temperature of the method of the present application and other comparative schemes. DETAILED DESCRIPTION
[0078] The present application will be further described in detail below in combination with the drawings and specific embodiments:
[0079] Example 1
[0080] As shown in Figure 1 , the present application provides a design flowchart of a smart home online energy management method combining prediction and learning ideas, including the following steps:
[0081] Under the premise of maintaining indoor user thermal comfort, the problem of minimizing the operation cost of the smart home energy management system is established, including the objective function, decision variables and constraint conditions, namely:
[0082]
[0083] wherein: is the expected operator; and are the energy cost and the depreciation cost of the battery energy storage system, respectively; is the maximum time slot number of the operation of the home energy management system; is the decision variable of the home energy management system at the moment, wherein: and respectively represent the charging and discharging power of the battery energy storage system at the moment, is the input power of the heating, ventilation and air conditioning system at the moment; in equation (2), and are the charging efficiency and the discharging efficiency of the battery energy storage system, respectively; in equation (3), , , are the minimum energy storage level, the maximum energy storage level and the energy storage level at the moment of the battery energy storage system, respectively; in equations (4)-(5), and are the maximum charging power and the maximum discharging power of the battery energy storage system, respectively; equation (6) indicates that the charging and discharging of the battery energy storage system cannot be carried out at the same time; in equation (7), , , are the inertia coefficient, the performance coefficient and the total heat transfer coefficient of the heating, ventilation and air conditioning system, respectively, represent the indoor temperature and the outdoor temperature at the moment, respectively; in equation (8), and represent the upper limit and the lower limit of the comfort temperature range, respectively; in equation (9), is the maximum input power of the heating, ventilation and air conditioning system; in equation (10), , and represent the amount of electricity purchased from the power grid, the photovoltaic power generation and the power demand of the rigid load at the moment, wherein: represents the amount of electricity sold to the power grid by the smart home energy system at the moment; represents the amount of electricity sold to the power grid by the smart home energy management system at the moment.At any given time, electricity is purchased from the grid; in equation (11), and They represent the first time. The purchase and sale prices of electricity at any given time; in equation (12), This is the depreciation unit cost coefficient for battery energy storage systems.
[0084] The problem of minimizing operating costs is remodeled as a Markov decision process related to the operation of a home energy management system, mainly including the states related to the operation of the smart home energy management system. Actions related to the control of battery energy storage systems and HVAC systems And the reward function related to the operating costs of the smart home energy management system and indoor thermal comfort. The specific expression is as follows:
[0085]
[0086] in: Indicates the first The time is relative to the time within a 24-hour period of the day; Indicates the battery energy storage system in the first The charging and discharging power at any given time, if but . Weighting coefficients used to measure the importance of operating costs and deviations from indoor thermal comfort. This indicates the penalty for deviation from indoor thermal comfort, i.e. Then, an agent is constructed to learn an implicit world model related to the operation of the home energy management system, comprising six neural networks, namely an encoder network. A dynamic network A reward network A policy network ,one Value function network set and a target Value function network set. The above neural networks all adopt a deep neural network structure, consisting of an input layer, multiple hidden layers, and an output layer. Specifically, the encoder network... Its input is the state observation value of the smart home energy management system. The output is a low-dimensional latent state code. The dynamic network Its input is the current potential state. Final execution action The output is the predicted potential state for the next time step. The reward network Its input is the current potential state. , final execution action , output the predicted immediate reward ; the policy network , input the current latent state , output the action mean and log standard deviation conforming to the Gaussian distribution, used for sampling to generate the preliminary action ; the value function network set is composed of multiple independent multi-layer perceptrons, and the input of each multi-layer perceptron is the current latent state , final execution action , output the state-action value function value. The above neural network constitutes an implicit world model related to the operation of the smart home energy management system, and its specific formula is as follows:
[0087]
[0088] The environment model of the agent interaction for building and learning the implicit world model of the home energy management system includes a battery energy storage system dynamic model and an indoor temperature dynamic model. The above models are constructed by using a white box modeling method. The data required for running the environment model includes photovoltaic power generation power, smart home rigid load, outdoor temperature, and real-time electricity price. The above data are obtained from the historical operation data of the smart home energy management system.
[0089] (1) According to the constructed environment model, an agent training framework based on the combination of model predictive control and deep reinforcement learning is established to train the agent to obtain an implicit world model related to the operation of the home energy management system. The agent training framework based on the combination of model predictive control and deep reinforcement learning specifically includes the following steps:
[0090] First, the neural network parameters of the agent, the experience replay pool , and the interaction environment model are initialized. Then, the number of time slots for the current agent to interact with the environment model is set . Finally, the preset iteration steps are repeatedly executed until the termination condition is met, and the iteration steps include:
[0091] At , it is determined whether the number of time slots for the current agent to interact with the environment model reaches the preset maximum random sampling action generation time slot number . If the current time slot number , the action is obtained by random sampling; otherwise, the agent generates the final execution action based on the current environment observation value and using the action planning function designed in combination with the model predictive control and deep reinforcement learning framework.The action planning function based on the model predictive control and deep reinforcement learning framework generates the final execution action The specific steps are as follows: first, the encoder network in the intelligent home energy management system hidden world model reads the state observation value at the moment and maps it to a low-dimensional latent state vector Second, the intelligent home energy management system hidden world model calls the hybrid policy to generate a set of candidate action sequences, each with a preset length of Specifically, the policy network of the intelligent home energy management system hidden world model first reads the latent state vector and performs step forward deduction to generate candidate action sequences. Then, based on the Gaussian distribution , generate candidate action sequences, where: and represent the mean and standard deviation in Gaussian distribution sampling, respectively. Then, based on the iterative model predictive path integral optimization algorithm, the candidate action sequences obtained are continuously optimized, and the specific steps include: in each iteration, except for the action sequences generated by the policy network in the home energy management system hidden world model , the remaining sequences obtained based on the random policy are randomly sampled. For each candidate action sequence, use the dynamic network and the reward network in the hidden world model to execute the candidate action sequence from the current latent state to perform step forward simulation and calculate the cumulative expected return of the simulation trajectory, the specific formula is as follows:
[0092]
[0093] Where: is the discount factor; is the mean of multiple independent multilayer perceptrons in the value function network set. According to the cumulative expected return of the candidate action sequences calculated in step , the optimal is selected from all candidate action sequences.These sequences constitute an elite action sequence set. Then, based on the obtained elite action sequence set, calculations are performed. The weighted average and weighted standard deviation are calculated, and the mean of the Gaussian distribution is updated using this result. and standard deviation Finally, repeat the preset number of iterations. The final iteration's action sequence output step includes: calculating the cumulative expected return based on the final elite action sequence set obtained from the previous step. A sampling weight is set for each elite action sequence, and an optimal action sequence is randomly selected from them using a differentiable discrete sampling method, and its first time step is extracted. action As the final output at the current decision-making moment Finally, in Superimposed from The sampled Gaussian noise, where: For the first time step The standard deviation.
[0094] (2) The agent will take action Execute within the constructed interactive environment to obtain new state observations. and reward function Next, the experience tuples generated from this interaction ( , , , Store in the experience replay pool middle.
[0095] (3) Determine the current interaction time slot Whether the update conditions for each neural network model in the agent have been met. If the update conditions have been met, then a small batch of data is sampled from the experience replay pool and the agent is updated based on the neural network update function that combines model predictive control and deep reinforcement learning. The specific steps include: First, in the experience replay pool... Random selection Each starting sampling time index is extracted and its length is [value missing]. First, obtain small batches of agent update data from continuous time segments. Second, sample reward values from the experience replay pool. and the observation at the next moment via encoder network Transform into the potential state of the next time step And input into the policy network. and target Value function network set for computation target value The specific calculation formula is as follows:
[0096]
[0097] in: For the goal The output of the value function network set is the minimum of multiple independent multilayer perceptrons. Then, the initial observations are... Encode as the initial potential state ; Utilize the sampled actions and initial potential state In dynamic space, perform a length of Multi-step evolution, through dynamic networks Generate multiple predicted next-time potential states And calculate its relationship with the true encoded latent state. Consistency loss between The specific calculation formula is as follows:
[0098]
[0099] in: As the loss weighting factor, This is the function to stop the gradient operation. Simultaneously, a length of... can be obtained. Predicted latent state sequence According to the obtained previous Potential states and the actions obtained from sampling Through reward networks and Value networks predict rewards respectively and TD error value The soft cross-entropy loss function is used to calculate the predicted reward. With real rewards The reward loss between, and Value and TD target value The value loss between them is calculated by weighting and summing all the above losses according to preset weighting coefficients to obtain the total loss function. The specific calculation formula is as follows:
[0100]
[0101] in: for Value function network parameters, Let cross-entropy be the loss function. For empirical tuples, The consistency loss coefficient, As the reward loss coefficient, is the value loss coefficient. According to the total loss function obtained, the gradient is calculated by the back propagation algorithm, and the parameters of all networks in the world model except the strategy network are optimized. The latent state sequence is input, the strategy loss is calculated by maximizing the value and the policy entropy , and the parameters of the strategy network are independently optimized, and the specific calculation formula is as follows:
[0102] (32)
[0103] Wherein: and are the moving statistical adjustment coefficients, which are used to automatically balance the relative size of the value and the entropy term, is the mean value of multiple independent multilayer perceptrons in the value function network. The target value network parameters are updated based on the soft update method, and the specific formula is as follows: (33)
[0104] Wherein: is the soft update coefficient.
[0105] (4) Let , and determine whether the current interaction time
[0106] reaches the preset maximum interaction time slot number . If the current time , jump to step (1) until ; if the current time , terminate the training, and obtain the implicit world model related to the operation of the smart home energy management system. (5) After training, decision-making is carried out based on the obtained implicit world model, and the decision-making is applied to the actual smart home system. This process only needs to call the encoder network and the strategy network
[0107] in the implicit world model of the home energy management system, that is:
[0108]
[0109] In order to show the effectiveness of the method of the present application, four comparative schemes are introduced.
[0110] The comparative scheme one adopts a rule-based home energy management strategy. Specifically, the comparative scheme one first controls the indoor thermal comfort based on a traditional heating, ventilation and air conditioning system on / off method. Taking the refrigeration mode as an example, when the indoor temperature is higher than the set upper limit of the indoor comfort temperature, the heating, ventilation and air conditioning system is turned on; when the indoor temperature is lower than the set lower limit of the indoor comfort temperature, the heating, ventilation and air conditioning system is turned off; and when the indoor temperature is within the indoor comfort temperature range, the current heating, ventilation and air conditioning system on-off state is maintained. Then, the current charging and discharging power of the battery energy storage system is solved based on the energy balance constraint of the home energy management system.
[0111] The comparative scheme two is a home energy management strategy adopting a traditional model predictive control algorithm.
[0112] The comparative scheme three adopts a traditional deep reinforcement learning method based on a deep deterministic policy gradient algorithm to optimize the operation of the smart home.
[0113] The comparative scheme four adopts a model-based deep reinforcement learning method to optimize the operation of the smart home. However, this method does not combine the model predictive control algorithm framework into the training process of the agent in the deep reinforcement learning algorithm framework.
[0114] Figures 2-4 is a comparison diagram of the method of the present application and the above four comparative schemes. As Figure 2 indicated, it is a comparison of the reward function training curves of the method of the present application and the comparative scheme three and the comparative scheme four. Specifically, compared with the comparative scheme three and the comparative scheme four, the method of the present application can converge to a higher reward value in the training process, and has better convergence.
[0115] As Figure 3 indicated, it is a comparison diagram of the operation cost of the method of the present application and the comparative schemes. From Figure 3 it can be seen that the method of the present application has achieved lower operation cost compared with all the comparative schemes. Specifically, compared with the comparative schemes one to four, the method of the present application can reduce the operation cost of the smart home by 18.05%, 16.11%, 12.41% and 11.10% respectively, and has better economic benefits.
[0116] As Figure 4 indicated, it is a comparison diagram of the indoor comfort temperature deviation of the method of the present application and the comparative schemes. From Figure 4 it can be seen that the method of the present application has achieved lower indoor comfort temperature deviation compared with all the comparative schemes. Specifically, compared with the comparative schemes one to four, the method of the present application can reduce the indoor comfort temperature deviation by 98.35%, 96.22%, 69.72% and 21.43% respectively, and can provide a more comfortable indoor environment.
[0117] The above merely describes the preferred embodiments of the present application, but does not constitute any other form of limitation to the present application, and any modification or equivalent change made according to the technical essence of the present application still falls within the scope of the present application.
Claims
1. A smart home online energy management method fusing the ideas of prediction and learning, characterized in that: Includes the following steps: (1) To minimize the operating cost of a smart home energy management system while maintaining the thermal comfort of indoor users, including the objective function, decision variables and constraints; (2) The problem of minimizing operating costs is remodeled as a Markov decision process related to the operation of the home energy management system, and an agent is constructed to learn the implicit world model related to the operation of the home energy management system. The agent in step (2) comprises six neural networks, i.e., an encoder network , a dynamics network , a reward network , a policy network , a value function network set and a target value function network set, all of which are deep neural networks comprising an input layer, multiple hidden layers and an output layer. In particular, the encoder network , whose input is the state observation of the smart home energy management system , and whose output is a low-dimensional latent state code ; the dynamic network , whose input is the current latent state , the final execution action , and whose output is the predicted next time latent state ; the reward network , whose input is the current latent state , the final execution action , and whose output is the predicted immediate reward ; the policy network , whose input is the current latent state , and whose output is the action mean and log standard deviation conforming to the Gaussian distribution, which are used for sampling to generate the preliminary action ; the value function network set is composed of multiple independent multi-layer perceptron integrations, the input of each multi-layer perceptron is the current latent state , the final execution action , and the output is the state-action value function value, and the above neural network constitutes an implicit world model related to the operation of the smart home energy management system, and its specific formula is as follows: ; (3) Construct an environment model for agent interaction in the implicit world model of the family energy management system; The environmental model that interacts with the agent learning the implicit world model of the home energy management system in step (3) includes a dynamic model of the battery energy storage system and a dynamic model of indoor temperature. The above models are constructed using a white-box modeling approach. The data required for the operation of the environmental model includes photovoltaic power generation, smart home rigid load, outdoor temperature and real-time electricity price. All of the above data are obtained from the historical operation data of the smart home energy management system. (4) Based on the constructed environment model, establish an agent training framework based on the combination of model predictive control and deep reinforcement learning, and train the agent to obtain the implicit world model related to the operation of the home energy management system. The agent training framework based on the combination of model predictive control and deep reinforcement learning in step (4) specifically includes the following steps: First, initialize the neural network parameters of the agent, the experience replay pool and the interaction environment model, then set the number of time slots for the current agent to interact with the environment model , and finally, repeat the preset iteration steps until the termination condition is met. The iterative steps in step (4) include: S51、in whether the number of time slots in which the current agent interacts with the environment model reaches a preset maximum number of time slots for generating actions by random sampling , if the current number of time slots , then the action is obtained by random sampling ; otherwise, the agent generates a final execution action based on the current environment observation value and uses an action planning function designed in combination with a model predictive control and deep reinforcement learning framework ; The action planning function designed in step S51 based on the model predictive control and deep reinforcement learning framework is specifically designed as follows: S61、First, the encoder network in the smart home energy management system implicit world model read state observation at time and map it to a low-dimensional latent state vector ; S62. Secondly, the implicit world model of the smart home energy management system calls a hybrid strategy to generate a model containing... A set of candidate action sequences, each sequence having a preset length of [length missing]. Specifically, the policy network of the implicit world model of the smart home energy management system. First, read the potential state vector. and carry out Step forward deduction, generate A candidate action sequence, then, based on a Gaussian distribution. Random sampling generation There are candidate action sequences, where: and Let represent the mean and standard deviation in a Gaussian distribution sample, respectively; S63. Then, the path integral optimization algorithm based on the iterative model continues to optimize the candidate action sequence obtained in step S62; candidate action sequences; S64、Finally, the output step of the final iteration of the action sequence is performed after the optimization of the candidate action sequence is completed. S64、Finally, the output step of the final iteration of the action sequence is performed after the optimization of the candidate action sequence is completed. S52, the agent will perform an action in the constructed interaction environment, obtaining a new state observation value and a reward function , then, storing the experience tuple generated by this interaction , , , into the experience replay pool ; S53、determine the current interaction time slot whether the update condition of each neural network model in the agent is reached, if the update condition is reached, a small batch of data is sampled from the experience replay pool, and the agent is updated based on the neural network update function combined with model predictive control and deep reinforcement learning; S54, let and determine whether the current interaction time reaches the preset maximum number of interaction time slots , if the current time , jump to step S51, until ; if the current time , terminate the training, and obtain the implicit world model related to the operation of the smart home energy management system; (5) After training, decisions are made based on the obtained implicit world model and applied to the actual smart home system; In step (5), decisions are made based on the obtained implicit world model and applied to the actual smart home system. This process only requires calling the encoder network in the implicit world model of the home energy management system. and policy network ,Right now .
2. The smart home online energy management method integrating prediction and learning concepts as described in claim 1, characterized in that: The problem of minimizing the operating cost of the smart home energy management system established in step (1) includes the objective function, decision variables, and constraints, namely: ; in: For expectation operators; and These are energy costs and depreciation costs of battery storage systems, respectively. This represents the maximum number of time slots that the home energy management system can operate on. For the first The decision variables for the home energy management system at any given time are: and These represent the battery energy storage system in the first... The charging and discharging power at any given time For HVAC systems in the first The input power at time t; in equation (2), and These are the charging efficiency and discharging efficiency of the battery energy storage system, respectively; in equation (3), , , These represent the minimum energy storage level, maximum energy storage level, and the first energy storage level of the battery energy storage system, respectively. The energy storage level at each moment; in equations (4)-(5), and These are the maximum charging power and maximum discharging power of the battery energy storage system, respectively; Equation (6) indicates that charging and discharging of the battery energy storage system cannot occur simultaneously; in Equation (7), , , These are the inertia coefficient, performance coefficient, and overall thermal conductivity of the HVAC system. They represent The indoor and outdoor temperatures at any given time; in equation (8), and represent the upper and lower limits of the comfortable temperature range, respectively; in equation (9), The maximum input power of the HVAC system; in equation (10), , and They represent the first time. The electricity purchased from the grid at all times, the photovoltaic power generation capacity, and the electricity demand of rigid loads, including: Indicates that the smart home energy system is in the first At any given moment, electricity is sold to the grid; The smart home energy management system indicates that in the first... At any given time, electricity is purchased from the grid; in equation (11), and They represent the first time. The purchase and sale prices of electricity at any given time; in equation (12), This is the depreciation unit cost coefficient for battery energy storage systems.
3. The smart home online energy management method integrating prediction and learning concepts as described in claim 2, characterized in that: The Markov decision process established in step (2) includes states related to the operation of the smart home energy management system. Actions related to the control of battery energy storage systems and HVAC systems And the reward function related to the operating costs of the smart home energy management system and indoor thermal comfort. The specific expression is as follows: ; in: Indicates the first The time is relative to the time within a 24-hour period of the day; Indicates the battery energy storage system in the first The charging and discharging power at any given time, if but , To measure the weighting of operating costs and deviations from indoor thermal comfort, This indicates the penalty for deviation from indoor thermal comfort, i.e. .
4. The smart home online energy management method integrating prediction and learning concepts as described in claim 3, characterized in that, In step S63, the path integral optimization algorithm based on the iterative model prediction is continuously optimized. The specific steps for each candidate action sequence are as follows: S71. First, in each iteration, in addition to the policy network in the implicit world model of the home energy management system... generated One action sequence, the rest obtained based on a random strategy. Each sequence is resampled randomly, and for each candidate action sequence, a dynamic network from the hidden world model is used. and reward network From the current potential state Begin executing the candidate action sequence to proceed. Perform a forward simulation and calculate the cumulative expected return of the simulated trajectory. The specific formula is as follows: ; in: Given a sequence of potential states, Preset length for each sequence, Discount factor; for The mean of multiple independent multilayer perceptrons in a set of value function networks; S72, Based on the calculation obtained in step S71 The cumulative expected return for each candidate action sequence Select the optimal action sequence from all candidate action sequences. These sequences constitute an elite action sequence set; S73. Based on the elite action sequence set obtained in step S72, calculate... The weighted average and weighted standard deviation are calculated, and the mean of the Gaussian distribution described in step S62 is updated with this result. and standard deviation ; S74. Repeat the preset number of iterations. The final iteration of the action sequence outputs the steps.
5. The smart home online energy management method integrating prediction and learning concepts as described in claim 4, characterized in that, The process is completed in step S64. The specific steps for outputting the final action sequence after optimizing the candidate action sequences are as follows: S81. The cumulative expected return based on the final elite action sequence set obtained in step S64. Set the sampling weight for each elite action sequence, and randomly select an optimal action sequence from them using a differentiable discrete sampling method; S82. Based on the optimal action sequence obtained from sampling in step S81, extract its first time step. action As the final output at the current decision-making moment ; S83, in the action Superimposed from The sampled Gaussian noise, where: For the first time step The standard deviation.
6. The smart home online energy management method integrating prediction and learning concepts as described in claim 5, characterized in that, The specific steps in step S53 of sampling small batches of data from the experience replay pool and updating the agent based on a neural network update function that combines model predictive control and deep reinforcement learning are as follows: S91. First, in the experience replay pool Random selection Each starting sampling time index is extracted and its length is [value missing]. From continuous time segments, obtain small batches of agent update data; S92. Secondly, the reward value sampled from the experience replay pool. and the observation at the next moment via encoder network Transform into the potential state of the next time step And input into the policy network. and target Value function network set for computation target value The specific calculation formula is as follows: ; in: For the goal The output of the value function network set is the minimum of at least two independent multilayer perceptrons; S93. Then, the initial observation values Encode as the initial potential state ; Utilize the sampled actions and initial potential state In dynamic space, perform a length of Multi-step evolution, through dynamic networks Generate multiple predicted next-time potential states And calculate its relationship with the true encoded latent state. Consistency loss between The specific calculation formula is as follows: ; in: As the loss weighting factor, To stop the gradient operation function, and simultaneously obtain a length of Predicted latent state sequence , This represents the predicted potential state at t=0 for the next time step. S94, Based on the previous information obtained in step S93 Potential states and the actions obtained from sampling Through reward networks and Value networks predict rewards respectively and TD error value The soft cross-entropy loss function is used to calculate the predicted reward. With real rewards The reward loss between, and Value and TD target value The value loss between them is calculated by weighting and summing all the above losses according to preset weighting coefficients to obtain the total loss function. The specific calculation formula is as follows: ; in: for Value function network parameters, Let cross-entropy be the loss function. For empirical tuples, The consistency loss coefficient, As the reward loss coefficient, This is the value loss coefficient; S95. Based on the total loss function obtained in step S94, calculate the gradient using the backpropagation algorithm, and apply this gradient to the world model excluding the policy network. Optimize the parameters of all networks; S96, Sequence of potential states As input, by maximizing Value and policy entropy Calculate the policy loss and apply it to the policy network. parameters Independent optimization is performed, and the specific calculation formula is as follows: (32) in: and This is an adjustment factor for moving statistics, used for automatic balancing. Value and entropy term The relative size, for The mean of multiple independent multilayer perceptrons in a value function network; S97. Update the target based on soft update method Value network parameters The specific formula is as follows: (33) in: This is the soft update coefficient.
Citation Information
Patent Citations
Smart home energy management method and system based on deep reinforcement learning
CN110458443A
Smart home energy management method and system based on personalized federal reinforcement learning
CN119903767A
Intelligent household energy management system prediction and decision integrated scheduling method based on deep reinforcement learning
CN116227883A
Household energy demand response optimization method and system based on deep reinforcement learning
CN117057553A