Residential energy storage system control method based on user behavior perception
By combining user behavior perception and reinforcement learning, Markov decision-making process is constructed, and the agent is trained to optimize the control strategy of energy storage systems and flexible resources, solving the problems of energy waste and user impact of real-time scheduling of energy storage systems in residential buildings, achieving efficient utilization and economic improvement.
Patent Information
- Application Number
- CN202510607041.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to effectively utilize energy storage systems and flexible resources in residential areas for real-time scheduling, ignoring user behavior leads to energy waste and affecting user experience.
Through the combination of user behavior perception and reinforcement learning, Markov decision-making process is constructed, and the agent is trained to optimize the control strategies of energy storage systems and flexible resources, including energy storage systems, HVACs, controllable loads and modeling of electric vehicles, and real-time scheduling is used for sensor data.
It realizes efficient utilization of residential energy storage systems, reduces energy waste, improves the potential of flexible resources to participate in distribution network scheduling, and reduces the impact on users, improving the economicality of system operation and user satisfaction.
Smart Images

Figure CN120511652A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electrical engineering, and more specifically, relates to a residential energy storage system control method based on user behavior perception. Background Art
[0002] With the growth of energy consumption in my country, energy-saving technologies and sustainable development have become topics of increasing concern. The United Nations Environment Programme pointed out in its 2022 Annual Status of the Building Industry Report that in 2021, residential energy consumption in operation accounted for 21% of the world's total final energy consumption and 17% of global carbon emissions. After the residential energy storage system and control system are equipped, the program can be used to manage the demand side, improve the flexibility of the residential building, and reasonably distribute energy, thereby reducing the building's energy consumption level. Early studies usually formulated the residential operation optimization problem as a sequential optimization problem, using mathematical methods or heuristic algorithms for optimization, such as linear programming, convex optimization, genetic algorithms, etc. However, these methods require some understanding of the residential environment and are difficult to apply to different residential scenarios. In addition, the proposed methods are only applicable to predictive control and are difficult to apply to real-time scheduling scenarios.
[0003] With the rapid development of artificial intelligence (AI), reinforcement learning (RL) has emerged as a method for solving optimal strategies. Compared to purely reactive heuristic control, RL can adjust control strategies based on changing objectives and learn optimal control actions through interaction with the environment without requiring any model development. Furthermore, RL is suitable for multi-objective cost function planning over long timescales and real-time scheduling, making it highly suitable for residential energy storage system control.
[0004] With the application of reinforcement learning in this scenario, questions about the impact of user behavior on residential operations have gradually emerged. User activity is undoubtedly one of the main causes of energy consumption, so ignoring user activity can lead to potential energy waste. Existing literature often ignores user activity within the home when applying control strategies to the environment. In addition, existing research on flexible resources within the home, such as energy storage systems, temperature control loads, and electric vehicles, often assumes that these types of loads are unconditionally involved in distribution network scheduling during demand response. However, existing research has shown that the use of these distributed energy resources can affect actual user usage. How to propose an effective strategy to reduce potential residential energy waste while leveraging the potential of residential energy storage systems and flexible resources to participate in the economic dispatch of the distribution network while minimizing the impact on users remains an open research issue. Summary of the Invention
[0005] In response to the defects of the existing technology and the need for improvement, the present invention provides a residential energy storage system control method based on user behavior perception, which aims to achieve efficient utilization of the residential energy storage system, solve the potential energy waste in the daily operation of the residence, and give full play to the potential of the residence to flexibly participate in the economic dispatch of the distribution network, while reducing the impact of this process on users.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions, including:
[0007] A residential energy storage system control method based on user behavior perception, comprising:
[0008] Step 1: Equivalent modeling of residential energy storage systems and flexibility resources
[0009] Model the energy storage system, HVAC system, controllable loads, and electric vehicles separately, defining their operating constraints and control parameters; construct power balance equations to ensure that the power interaction between the residence and the distribution grid meets real-time power balance;
[0010] Step 2: User behavior perception and scenario division
[0011] Data is collected through deployed power sensors and environmental sensors, distinguishing between Class I sensors and Class II sensors. Class I sensors are controlled only by user interaction, while Class II sensors are partially automatically controlled.
[0012] Identify user activities based on the correlation between main appliances, auxiliary appliances and environmental feedback;
[0013] Based on the user behavior perception results, residential operation is divided into scenario I with energy consumption-related activities and scenario II without related activities;
[0014] Step 3: Markov decision process modeling of residential operations
[0015] Construct Markov decision processes (MDPs) for scenario I and scenario II respectively;
[0016] The state space of scenario I includes the energy storage state of charge, time-of-use electricity price, electric vehicle power, and user comfort parameters. The action space is the combination of controllable loads, HVAC, and electric vehicle control coefficients. The reward function is designed based on electricity cost, user satisfaction, and grid interaction stability.
[0017] The state space of Scenario II includes energy storage power, electricity price deviation, and minimum controllable load. The action space is the combination of control coefficients of the HVAC, electric vehicle, and energy storage system. The reward function is based on the grid peak-shaving subsidy and energy waste suppression design.
[0018] Step 4: Dynamic Control Based on Behavior Perception and Reinforcement Learning
[0019] The SoftActor-Critic (SAC) algorithm is used to train the agents for scenario I and scenario II respectively;
[0020] The agent generates control actions through the policy network, optimizes the Q value through the critic network, and improves the exploration ability by combining entropy regularization;
[0021] Process sensor data in real time, call the intelligent agent of the corresponding scene according to the user behavior perception results, and dynamically adjust the operating parameters of the energy storage system, controllable load and electric vehicle.
[0022] Furthermore, the energy storage system is controlled by the charge and discharge power coefficient, and the constraint that charging and discharging cannot be performed simultaneously is satisfied;
[0023] The HVAC control method is achieved through the cooling / heating power coefficient to ensure that the indoor temperature is within the preset comfort range;
[0024] The control mode of controllable load is realized through power regulation coefficient and meets the minimum / maximum power limit;
[0025] The charging power of electric vehicles is dynamically adjusted through the efficiency coefficient and the current power.
[0026] Furthermore, in step 1, the energy storage system, HVAC system, controllable load and electric vehicle are modeled and the power balance equation is constructed, including:
[0027] (1.1) Energy storage system modeling
[0028] The energy storage charging and discharging process is:
[0029]
[0030] Where, and is the electric energy stored at the current moment, η es,c and η es,d is the charging efficiency and discharging efficiency of energy storage, and is the charging and discharging power of the energy storage;
[0031] In addition, energy storage must also meet the following constraints: power constraint, the energy storage charge and discharge power cannot exceed the upper limit, and charging and discharging cannot be carried out simultaneously; capacity constraint, the energy storage charge and discharge process must be carried out within the discharge depth; the energy storage control method is:
[0032]
[0033] In the formula, coefficient a es Used to control the charging and discharging of energy storage systems, and The maximum charging and discharging power for energy storage;
[0034] (1.2) HVAC system modeling:
[0035] The HVAC working process is:
[0036]
[0037] Where, and is the indoor temperature at the current moment and the previous moment, R m and C m are the thermal resistance and heat capacity of the room, is the active power of the HVAC, is the current outdoor temperature. The role of HVAC is to keep the indoor temperature within a comfortable range as much as possible:
[0038]
[0039] Where, T min and T max are the lower and upper limits of the indoor comfort temperature, respectively. In addition, when the HVAC is used as a temperature control load, it must operate within this range. The control method of the HVAC is:
[0040]
[0041] In the formula, coefficient a hvac Used to control HVAC cooling or heating, and The maximum heating and cooling power of the HVAC;
[0042] (1.3) Controllable load modeling
[0043] Controllable load refers to the load whose power can be adjusted within a certain range. Its operation and control methods are as follows:
[0044]
[0045] Where, and is the current minimum power and maximum power of the controllable load, is the control coefficient of the controllable load;
[0046] (1.4) Electric vehicle modeling
[0047] The charging and discharging process of electric vehicles is:
[0048]
[0049] Where, is the current power of the electric vehicle, η ev,c For electric vehicle charging efficiency, Charging power for electric vehicles;
[0050] (1.5) Power balance equation
[0051] Residential properties can sell excess electricity to the distribution grid or buy back electricity from the distribution grid, but the following power balance equation must be satisfied:
[0052]
[0053] Where, is the rooftop photovoltaic power generation power, For power sold to the distribution grid or bought back, is the power of the uncontrollable load.
[0054] Furthermore, user behavior perception includes detecting the startup status of major electrical appliances and judging the continuity of activities by combining auxiliary appliance data and environmental feedback;
[0055] The trigger condition for scenario I is the detection of appliance usage behavior that is directly related to energy consumption;
[0056] The triggering condition for scenario II is that no electrical appliance usage behavior directly related to energy consumption is detected and the load is in idle or automatic operation state.
[0057] Furthermore, the reward function of scenario I includes the electricity cost term, the user comfort term, and the electric vehicle power anxiety term;
[0058] The reward function of scenario II includes a grid peak-shaving subsidy item, an energy waste penalty item, and an electric vehicle power maintenance item.
[0059] Furthermore, step three specifically includes:
[0060] The goal of the agent in scenario I is to save residential energy expenditure by adjusting controllable loads, HVAC, and electric vehicle charging power while ensuring user satisfaction and taking into account dynamic electricity price changes;
[0061] (a) The state space consists of all the environmental variables that may affect the agent’s decision:
[0062]
[0063] Where t is the current time, A Boolean value indicating whether the electric vehicle can be charged. is the current state of charge of the electric vehicle, is the current state of charge of the energy storage system, and are the current purchase price and selling price of unit electricity respectively;
[0064] (b) The agent's actions are determined by controllable factors:
[0065]
[0066] Where a con , a hvac , a ev are the control coefficients of controllable loads, HVAC, and electric vehicles, respectively;
[0067] (c) The agent’s reward consists of the following components:
[0068]
[0069] Where, It describes the cost of purchasing energy from the distribution grid or the profit of selling energy to the distribution grid, and k1 is the reward coefficient;
[0070]
[0071] Where, It describes the user dissatisfaction caused by the decrease of controllable load. K2 is the reward coefficient, which is used to measure whether the user is willing to maintain the controllable load at a high level to ensure the user experience.
[0072]
[0073] Where, It describes the user dissatisfaction caused by the indoor temperature deviating from the comfort range. K3 is the reward coefficient, which is used to measure whether the user is willing to increase the HVAC power to ensure the user experience.
[0074]
[0075] Where, Describes the user's anxiety level about the current battery life of electric vehicles, is the basic state of charge level for normal use of electric vehicles, and k4 is the reward coefficient, which is used to measure the willingness to charge the electric vehicle to a higher level to alleviate user anxiety;
[0076] In scenario II, the goal of the intelligent agent is to shut down or reduce the relevant loads when there is no user behavior to avoid potential energy waste, balance the changes in energy prices in the market to reflect the peak and valley conditions of the system, and effectively leverage the potential of residential energy storage systems and flexibility resources to participate in system peak regulation;
[0077] (a) The state space consists of all the environmental variables that may affect the agent’s decision:
[0078]
[0079] Where t is the current time, is the current minimum value of the controllable load, is the power currently absorbed or released by the energy storage system, The difference between the current electricity price and the average price for the day;
[0080] (b) The agent's actions are determined by controllable factors:
[0081]
[0082] Where a hvac , a ev , a es are the control coefficients of HVAC, electric vehicles, and energy storage systems, respectively;
[0083] (c) The agent’s reward consists of the following components:
[0084]
[0085] Where, It describes the cost of the residential building to purchase energy from the distribution grid or the profit obtained by selling energy to the distribution grid, and b1 is the incentive coefficient;
[0086]
[0087]
[0088] Where, The subsidy for residential energy storage systems and flexible resources participating in distribution network scheduling to reduce peak loads is described. When distributed energy resources do not participate in peak load reduction and filling, no subsidy is obtained. is the day-ahead electricity price, b2 is the incentive coefficient;
[0089]
[0090] Where, It describes the user's anxiety level about the current electric vehicle power level, and b3 is the reward coefficient, which is used to measure whether the user is willing to charge the electric vehicle to a higher level to alleviate the user's anxiety.
[0091] Furthermore, step four specifically includes: setting up two intelligent agents for scenarios I and II respectively to execute the corresponding Markov decision process; during the training process, setting up an independent environment for each intelligent agent, and using the SAC algorithm to train the intelligent agent; after obtaining an intelligent agent that can be used for real-time scheduling, processing sensor data and user behavior perception, thereby selecting the corresponding intelligent agent and executing the corresponding energy storage system control strategy.
[0092] Compared with the prior art, the above technical solution proposed by the present invention can achieve the following effects:
[0093] 1. A method for perceiving user behavior in a residence through sensor data is proposed, and the environment is divided into scene I and scene II.
[0094] 2. Taking into account user satisfaction and system economics, the operation of residential energy storage systems and various flexibility resources is optimized by introducing intelligent agents with new states and rewards to adapt to different residential environments.
[0095] 3. A new flexibility resource scheduling scheme is proposed to schedule residential flexibility resources according to system price signals, thereby improving the economic efficiency of system operation while minimizing the impact on users. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 A flowchart of a residential energy storage system control method based on user behavior perception provided by an embodiment of the present invention;
[0097] Figure 2 A schematic diagram of a residential energy storage system control method based on user behavior perception provided by an embodiment of the present invention;
[0098] Figure 3 The reward change trend during the agent training process of the embodiment of the present invention;
[0099] Figure 4 The performance of the strategy proposed in the embodiment of the present invention on a certain day's data in the data set, including controllable loads, residential temperature control, energy storage systems, electric vehicles, and regulation based on dynamic prices;
[0100] Figure 5 The energy-saving performance of the strategy proposed in the embodiment of the present invention;
[0101] Figure 6 This is the performance of the strategy proposed in the embodiment of the present invention in the economic benefits of residential operation. DETAILED DESCRIPTION
[0102] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0103] like Figure 1 As shown, the present invention provides a residential energy storage system control method based on user behavior perception, which includes the following steps:
[0104] Equivalent modeling of residential energy storage systems and flexibility resources.
[0105] Analyze sensor data to perceive user behavior in the home.
[0106] A Markov decision process representation of residential operation is proposed, and the state, action and reward of the intelligent agent are proposed.
[0107] Determine a control strategy based on behavior perception and reinforcement learning, use this approach to train an intelligent agent, and perform real-time control of the residence.
[0108] Example
[0109] This example uses the smart home dataset provided by the U.S. Department of Energy's Office of Building Technologies as an example.
[0110] The house includes controllable loads, HVAC systems, uncontrollable loads, energy storage systems, rooftop photovoltaics and electric vehicles. Power sensors are deployed on most household appliances. It is also equipped with temperature sensors and water flow monitoring sensors with a data resolution of 1 minute.
[0111] like Figure 1 As shown, the main steps of this example include:
[0112] Step 1: Equivalent modeling of residential energy storage systems and flexibility resources.
[0113] (1.1) Energy storage system modeling. The energy storage charging and discharging process is
[0114]
[0115] Where, and is the electric energy stored at the current moment, η es,c and η es,d is the charging efficiency and discharging efficiency of energy storage, and is the charging and discharging power of the energy storage.
[0116] In addition, energy storage also needs to meet the following constraints: power constraint, the energy storage charging and discharging power cannot exceed the upper limit, and charging and discharging cannot be performed at the same time; capacity constraint, the energy storage charging and discharging process must be carried out within the discharge depth. The energy storage control method is:
[0117]
[0118] In the formula, coefficient a es Used to control the charging and discharging of energy storage systems, and The maximum charging and discharging power of energy storage.
[0119] (1.2) HVAC system modeling. The HVAC working process is
[0120]
[0121] Where, and is the indoor temperature at the current moment and the previous moment, R m and C m are the thermal resistance and heat capacity of the room, is the active power of the HVAC, is the current outdoor temperature. The role of HVAC is to keep the indoor temperature within a comfortable range as much as possible:
[0122]
[0123] Where, T min and T max These are the lower and upper limits of the indoor comfortable temperature, respectively. In addition, when the HVAC is used as a temperature control load, it must operate within this range. The control method of the HVAC is:
[0124]
[0125] In the formula, coefficient a hvac Used to control HVAC cooling or heating, and Maximum heating and cooling power for HVAC
[0126] (1.3) Controllable load modeling. A controllable load is a load whose power can be adjusted within a certain range. Its operation and control methods are as follows:
[0127]
[0128]
[0129] Where, and is the current minimum power and maximum power of the controllable load, is the control coefficient of the controllable load.
[0130] (1.4) Electric vehicle modeling. The charging and discharging process of an electric vehicle is:
[0131]
[0132] Where, is the current power of the electric vehicle, η ev,c For electric vehicle charging efficiency, Charging power for electric vehicles.
[0133] (1.5) Power balance equation. Residential buildings can sell excess electricity to the distribution grid or repurchase electricity from the distribution grid, but the following power balance equation must be satisfied:
[0134]
[0135] Where, is the rooftop photovoltaic power generation power, For power sold to the distribution grid or bought back, is the power of the uncontrollable load.
[0136] Step 2: User behavior perception method.
[0137] Combining the power sensor measurements with the feedback from the environmental sensors, the proposed strategy uses a hierarchical approach to identify the user's energy consumption-related behaviors within the home. The main steps include:
[0138] (2.1) Select user activities: exclude all activities that occur outside the home, select those user activities that are only related to residential energy consumption, and exclude all user activities performed using mobile power supplies.
[0139] (2.2) Sensor data processing: Based on the location of power sensors, they can be divided into Class I sensors, where the state of the deployed electrical appliances can only be changed through user interaction; and Class II sensors, where the state of the deployed electrical appliances can be changed by users but also have some automatic capabilities.
[0140] (2.3) User Behavior Mapping: This paper describes user behavior using primary appliances, auxiliary appliances, and environmental feedback. Primary appliances can only be used for a specific activity, or must be used for a specific activity; auxiliary appliances may be used for a specific activity; and environmental feedback describes changes in the environment caused by user behavior. Furthermore, a given appliance can be used for multiple activities, and multiple activities can occur simultaneously in the environment. Furthermore, environmental feedback is not a necessary condition for determining whether an activity is being performed.
[0141] (2.4) User behavior inference: The process of perceiving user behavior through sensors is as follows: if it is detected that the activation of a main appliance corresponds to a certain activity, the sensor data of the main appliance and auxiliary appliances, combined with environmental feedback, are used to determine whether the activity is ongoing. If a certain activity does not involve a main appliance, the sensor data of the auxiliary appliances are used to determine whether the activity is ongoing. In other cases, it is determined that no user activity has occurred.
[0142] Based on the results of user behavior perception in the residence, the environment can be divided into: Scenario I, user behavior related to energy consumption is detected; Scenario II, no user behavior related to energy consumption is detected.
[0143] Step 3: Markov decision process representation of residential operation.
[0144] Residential operation can be viewed as a sequential decision-making process on controllable variables. Therefore, it can be represented by a Markov decision process (MDP), which is typically represented by five parameters: S, A, P, R, and γ. In an MDP, the state space is the set of all possible states, typically denoted S; the action space is the set of all possible actions that the agent can take, typically denoted A. Based on its decision, the agent chooses an action to influence the environment. The transition probability describes the probability distribution of transitioning to a new state after taking an action in a specific state. It is typically denoted P(s'|s,a), where s represents the current state, a represents the action taken, and s' represents the new state. The reward function R(s,a,s') defines the immediate reward the agent receives when transitioning to a new state after taking a specific action in a specific state. The discount factor γ∈[0,1] represents the importance the agent places on future rewards.
[0145] (3.1) Scenario I
[0146] The goal of the intelligent agent in scenario I is to save residential energy expenditure by adjusting controllable loads, HVAC, and electric vehicle charging power while ensuring user satisfaction and taking into account dynamic electricity price changes.
[0147] (a) The state space consists of all the environmental variables that may affect the agent’s decision:
[0148]
[0149] Where t is the current time, A Boolean value indicating whether the electric vehicle can be charged. is the current state of charge of the electric vehicle, is the current state of charge of the energy storage system, and are the current purchase price and selling price of unit electricity respectively.
[0150] (b) The agent's actions are determined by controllable factors:
[0151]
[0152] Where a con , a hvac , a ev are the control coefficients of controllable loads, HVAC, and electric vehicles, respectively.
[0153] (c) The agent’s reward consists of the following components:
[0154]
[0155] Where, It describes the cost of purchasing energy from the distribution grid or the profit of selling energy to the distribution grid, and k1 is the reward coefficient.
[0156]
[0157] Where, It describes the user dissatisfaction caused by the decrease of controllable load. K2 is the reward coefficient, which is used to measure whether the user is willing to maintain the controllable load at a high level to ensure the user experience.
[0158]
[0159] Where, It describes the user dissatisfaction caused by the indoor temperature deviating from the comfort range. K3 is the reward coefficient, which is used to measure whether the user is willing to increase the HVAC power to ensure the user experience.
[0160]
[0161] Where, Describes the user's anxiety level about the current battery life of electric vehicles, is the basic state of charge level for normal use of electric vehicles, and k4 is the reward coefficient, which is used to measure whether the user is willing to charge the electric vehicle to a higher level to alleviate the user's anxiety.
[0162] (3.2) Scenario II
[0163] The goal of the intelligent agent in scenario II is to shut down or reduce the relevant loads when there is no user behavior to avoid potential waste of electricity, balance the changes in energy prices in the market to reflect the peak and valley conditions of the system, and effectively tap the potential of residential energy storage systems and flexibility resources to participate in system peak regulation.
[0164] (a) The state space consists of all the environmental variables that may affect the agent’s decision:
[0165]
[0166] Where t is the current time, is the current minimum value of the controllable load, is the power currently absorbed or released by the energy storage system, It is the difference between the current electricity price and the average price for the day.
[0167] (b) The agent's actions are determined by controllable factors:
[0168]
[0169] Where a hvac , a ev , a es are the control coefficients of HVAC, electric vehicles, and energy storage systems, respectively.
[0170] (c) The agent’s reward consists of the following components:
[0171]
[0172] Where, It describes the cost of purchasing energy from the distribution network or the profit of selling energy to the distribution network, and b1 is the reward coefficient.
[0173]
[0174] Where, The paper describes the subsidies that residential energy storage systems and flexible resources receive when participating in distribution network dispatch to reduce peak loads. Obviously, when distributed energy resources do not participate in reducing peak loads, they will not receive subsidies. is the day-ahead electricity price, b2 is the incentive coefficient, which varies due to changes in different policies.
[0175]
[0176] Where, It describes the user's anxiety level about the current electric vehicle power level, and b3 is the reward coefficient, which is used to measure whether the user is willing to charge the electric vehicle to a higher level to alleviate the user's anxiety.
[0177] Step 4: Control strategy based on behavior perception and reinforcement learning.
[0178] (4.1) SAC algorithm
[0179] The SAC algorithm is an extension of the actor-critic algorithm, which is used to solve reinforcement learning problems in continuous action spaces. The agent initializes the policy network, critic network, target Q network and replay buffer of the SAC algorithm, using the current policy π φ Select action a t , get feedback by interacting with the environment, that is, the next moment state s t+1 , reward r t and the termination symbol d t , and the experience (s t ,a t ,r t ,s t+1 ,d t ) is stored in the replay buffer; then, the agent samples a batch of data from the replay buffer and φ Sample the next action a in (a|s) t+1 , calculate its logarithmic probability logπ φ (a t+1 |s t+1 ), and calculate the Q value of the target:
[0180]
[0181] The agent then updates the critic Q network using the mean squared error and Parameters:
[0182]
[0183] The agent minimizes the critic network output and As the goal, update the policy network π φ Parameters of (a|s):
[0184]
[0185] Update the parameters of the target Q network using the soft method:
[0186] θ′ i ←τθ i +(1-τ)θ′ i (26)
[0187] Where τ is the soft update coefficient.
[0188] Finally, the agent automatically adjusts the entropy coefficient α using the gradient descent method, setting the target entropy value to And update the entropy value α through the following loss function:
[0189]
[0190] (4.2) Control strategy based on behavior perception and reinforcement learning
[0191] For scenarios I and II, two agents are set up to execute the corresponding Markov decision process. During the training process, an independent environment is set up for each agent, and the SAC algorithm in (4.1) is used to train the agent. After obtaining the agent that can be used for real-time scheduling, the sensor data is processed to perceive user behavior, thereby selecting the corresponding agent and executing the corresponding energy storage system control strategy. The detailed process of the algorithm is as follows: Figure 3 shown.
[0192] The independent learning algorithm and hyperparameters for each energy storage system are the same. The hyperparameters used in this case training are shown in Table 1:
[0193] Table 1 Training hyperparameters
[0194] parameter Numerical Learning rate of the actor network <![CDATA[3×10 -4 ]]> Learning rate of the critic network <![CDATA[3×10 -4 ]]> The learning rate of α <![CDATA[3×10 -4 ]]> Total rounds 200 γ 0.99 Batch size 256 Playback buffer size 10000
[0195] The simulation was built using Python / Pytorch, and the detailed environment information is shown in Table 2. During the initial training phase, since the agent is still in the trial phase, the penalty for interacting with the environment is very large. Subsequently, the distributed controller gradually stores the better performance in the replay buffer and learns to improve the control strategy. After 1200 and 4000 training rounds, respectively, the performance indicators of the two agents gradually stabilized near the optimal value, as shown in Figure 2. Figure 3 shown.
[0196] Table 2 Detailed parameters of residential buildings
[0197]
[0198] To evaluate the transient stability performance of the proposed strategy in detail, the trained agent was used to simulate residential operation on a single day in July. Figure 4 The results show that user behavior perception can obtain the ongoing activities of users in the residence, which are represented by dark areas on the timeline. The proposed strategy can effectively control the use of energy storage systems, HVAC and controllable loads, shut down related loads when there is no user activity in the residence, and reduce potential energy waste. At the same time, when the current energy price is higher or lower than the average price, the residential energy storage system and flexibility resources are dispatched during peak or valley hours to participate in peak load regulation to obtain profits.
[0199] Before the control method disclosed in the present invention is adopted, if the residential energy storage system control strategy is not used in the residence, as shown in baseline 1, and if only the ordinary residential energy storage system control strategy is used without considering user behavior perception, as shown in baseline 2, the effects of different control strategies are evaluated over a period of one month. The results show that the strategy proposed in the present invention can effectively reduce energy consumption and improve the economic benefits of residential operation. Figure 5 and Figure 6 As shown, compared with no real-time control, the residential energy consumption was reduced by 23.9%, and the income from residential operation increased from 25.52€ to 53.8€. In the same environment, using only ordinary residential energy storage system control strategies, the energy expenditure savings were only 4.55%.
[0200] Those skilled in the art will understand that the foregoing descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art will be able to modify the technical solutions described in the foregoing examples or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention shall be included within the scope of protection of the invention.
Claims
1. A residential energy storage system control method based on user behavior perception, characterized in that: include: Step 1: Equivalent modeling of residential energy storage systems and flexibility resources Model the energy storage system, HVAC system, controllable loads, and electric vehicles separately, defining their operating constraints and control parameters; construct power balance equations to ensure that the power interaction between the residence and the distribution grid meets real-time power balance; Step 2: User behavior perception and scenario division Data is collected through deployed power sensors and environmental sensors, distinguishing between Class I sensors and Class II sensors. Class I sensors are controlled only by user interaction, while Class II sensors are partially automatically controlled. Identify user activities based on the correlation between main appliances, auxiliary appliances and environmental feedback; Based on the user behavior perception results, residential operation is divided into scenario I with energy consumption-related activities and scenario II without related activities; Step 3: Markov decision process modeling of residential operations Construct Markov decision processes (MDPs) for scenario I and scenario II respectively; The state space of scenario I includes the energy storage state of charge, time-of-use electricity price, electric vehicle power, and user comfort parameters. The action space is the combination of controllable loads, HVAC, and electric vehicle control coefficients. The reward function is designed based on electricity cost, user satisfaction, and grid interaction stability. The state space of Scenario II includes energy storage power, electricity price deviation, and minimum controllable load. The action space is the combination of control coefficients of the HVAC, electric vehicle, and energy storage system. The reward function is based on the grid peak-shaving subsidy and energy waste suppression design. Step 4: Dynamic Control Based on Behavior Perception and Reinforcement Learning The SoftActor-Critic (SAC) algorithm is used to train the agents for scenario I and scenario II respectively; The agent generates control actions through the policy network, optimizes the Q value through the critic network, and improves the exploration ability by combining entropy regularization; Process sensor data in real time, call the intelligent agent of the corresponding scene according to the user behavior perception results, and dynamically adjust the operating parameters of the energy storage system, controllable load and electric vehicle.
2. The method according to claim 1, characterized in that In step one: The energy storage system is controlled by the charge and discharge power coefficient, and the constraint of not being able to charge and discharge simultaneously is met; The HVAC control method is achieved through the cooling / heating power coefficient to ensure that the indoor temperature is within the preset comfort range; The control mode of controllable load is realized through power regulation coefficient and meets the minimum / maximum power limit; The charging power of electric vehicles is dynamically adjusted through the efficiency coefficient and the current power.
3. The method according to claim 2, characterized in that In step 1, the energy storage system, HVAC system, controllable load, and electric vehicle are modeled separately and the power balance equation is constructed, including: (1.1) Energy storage system modeling The energy storage charging and discharging process is: Where, and is the electric energy stored at the current moment, η es,c and η es,d is the charging efficiency and discharging efficiency of energy storage, and is the charging and discharging power of the energy storage; In addition, energy storage must also meet the following constraints: power constraint, the energy storage charge and discharge power cannot exceed the upper limit, and charging and discharging cannot be carried out simultaneously; capacity constraint, the energy storage charge and discharge process must be carried out within the discharge depth; the energy storage control method is: In the formula, coefficient a es Used to control the charging and discharging of energy storage systems, and The maximum charging and discharging power for energy storage; (1.2) HVAC system modeling: The HVAC working process is: Where, and is the indoor temperature at the current moment and the previous moment, R m and C m are the thermal resistance and heat capacity of the room, is the active power of the HVAC, is the current outdoor temperature. The role of HVAC is to keep the indoor temperature within a comfortable range as much as possible: T min ≤T t in ≤T max (4) Where, T min and T max are the lower and upper limits of the indoor comfort temperature, respectively. In addition, when the HVAC is used as a temperature control load, it must operate within this range. The control method of the HVAC is: In the formula, coefficient a hvac Used to control HVAC cooling or heating, and The maximum heating and cooling power of the HVAC; (1.3) Controllable load modeling Controllable load refers to the load whose power can be adjusted within a certain range. Its operation and control methods are as follows: Where, and is the current minimum power and maximum power of the controllable load, is the control coefficient of the controllable load; (1.4) Electric vehicle modeling The charging and discharging process of electric vehicles is: Where, is the current power of the electric vehicle, η ev,c For electric vehicle charging efficiency, Charging power for electric vehicles; (1.5) Power balance equation Residential properties can sell excess electricity to the distribution grid or buy back electricity from the distribution grid, but the following power balance equation must be satisfied: Where, is the rooftop photovoltaic power generation power, For power sold to the distribution grid or bought back, is the power of the uncontrollable load.
4. The method according to claim 1, wherein In step 2: User behavior perception includes detecting the startup status of major electrical appliances and judging activity continuity by combining auxiliary appliance data and environmental feedback; The trigger condition for scenario I is the detection of appliance usage behavior that is directly related to energy consumption; The triggering condition for scenario II is that no electrical appliance usage behavior directly related to energy consumption is detected and the load is in idle or automatic operation state.
5. The method according to claim 1, wherein In step three: The reward function of scenario I includes electricity cost term, user comfort term and electric vehicle power anxiety term; The reward function of scenario II includes a grid peak-shaving subsidy item, an energy waste penalty item, and an electric vehicle power maintenance item.
6. The method according to claim 3, characterized in that Step three specifically includes: The goal of the agent in scenario I is to save residential energy expenditure by adjusting controllable loads, HVAC, and electric vehicle charging power while ensuring user satisfaction and taking into account dynamic electricity price changes; (a) The state space consists of all the environmental variables that may affect the agent’s decision: Where t is the current time, A Boolean value indicating whether the electric vehicle can be charged. is the current state of charge of the electric vehicle, is the current state of charge of the energy storage system, and are the current purchase price and selling price of unit electricity respectively; (b) The agent's actions are determined by controllable factors: Where a con , a hvac , a ev are the control coefficients of controllable loads, HVAC, and electric vehicles, respectively; (c) The agent’s reward consists of the following components: Where, It describes the cost of purchasing energy from the distribution grid or the profit of selling energy to the distribution grid, and k1 is the reward coefficient; Where, It describes the user dissatisfaction caused by the decrease of controllable load. K2 is the reward coefficient, which is used to measure whether the user is willing to maintain the controllable load at a high level to ensure the user experience. r t temp =k3([T t in -T max ] + +[T min -T t in ] + ) (15) Where, It describes the user dissatisfaction caused by the indoor temperature deviating from the comfort range. K3 is the reward coefficient, which is used to measure whether the user is willing to increase the HVAC power to ensure the user experience. Where, Describes the user's anxiety level about the current battery life of electric vehicles, is the basic state of charge level for normal use of electric vehicles, and k4 is the reward coefficient, which is used to measure the willingness to charge the electric vehicle to a higher level to alleviate user anxiety; In scenario II, the goal of the intelligent agent is to shut down or reduce the relevant loads when there is no user behavior to avoid potential energy waste, balance the changes in energy prices in the market to reflect the peak and valley conditions of the system, and effectively leverage the potential of residential energy storage systems and flexibility resources to participate in system peak regulation; (a) The state space consists of all the environmental variables that may affect the agent’s decision: Where t is the current time, is the current minimum value of the controllable load, is the power currently absorbed or released by the energy storage system, The difference between the current electricity price and the average price for the day; (b) The agent's actions are determined by controllable factors: Where a hvac , a ev , a es are the control coefficients of HVAC, electric vehicles, and energy storage systems, respectively; (c) The agent’s reward consists of the following components: r t pft =b1r t home (19) Where, It describes the cost of the residential building to purchase energy from the distribution grid or the profit obtained by selling energy to the distribution grid, and b1 is the incentive coefficient; Where, The subsidy for residential energy storage systems and flexible resources participating in distribution network scheduling to reduce peak loads is described. When distributed energy resources do not participate in peak load reduction and filling, no subsidy is obtained. is the day-ahead electricity price, b2 is the incentive coefficient; Where, It describes the user's anxiety level about the current electric vehicle power level, and b3 is the reward coefficient, which is used to measure whether the user is willing to charge the electric vehicle to a higher level to alleviate the user's anxiety.
7. The method according to claim 1, characterized in that Step 4 specifically includes: setting up two intelligent agents for scenarios I and II respectively to execute the corresponding Markov decision process. During the training process, an independent environment is set up for each intelligent agent, and the SAC algorithm is used to train the intelligent agent. After obtaining an intelligent agent that can be used for real-time scheduling, user behavior perception is performed by processing sensor data, thereby selecting the corresponding intelligent agent and executing the corresponding energy storage system control strategy.