A building load cooperative regulation method considering user demand of V2B mode

By implementing the improved deep reinforcement learning algorithm in the patent, the charging and discharging strategy of electric vehicle batteries is optimized and user privacy is protected. This solves the dynamic and privacy issues in the collaborative control of building load in the V2B model, and realizes the efficient use of energy storage equipment and enhances user trust.

CN120016548BActive Publication Date: 2025-11-21HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510081810.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-11-21
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In the V2B model, the dynamic and random nature of building load demand and electric vehicle charging and discharging behavior makes it difficult for traditional control methods to cope, and existing deep reinforcement learning algorithms have failed to effectively protect user privacy, affecting the lifespan of energy storage devices and user acceptance.

Method used

An improved deep reinforcement learning algorithm is constructed to establish a mathematical model by simulating the power changes of building photovoltaics and rigid loads, optimize the charging and discharging strategy of electric vehicle batteries, and introduce a privacy protection mechanism to protect sensitive user data, extend the life of energy storage equipment, and enhance user trust.

Benefits of technology

It effectively extends the lifespan of energy storage devices, reduces battery replacement costs, increases user acceptance of the V2B model, reduces the risk of privacy leaks, and enhances user trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016548B_ABST
    Figure CN120016548B_ABST
Patent Text Reader

Abstract

The application discloses a kind of building load collaborative control methods considering V2B mode user demand, comprising: obtaining building photovoltaic, rigid load power model;The mathematical model of distributed generator is constructed;Electric vehicle battery energy storage mathematical model is constructed;Air conditioning load mathematical model is constructed;Establish power balance constraint equation;Based on power balance constraint equation, the cost model of building load collaborative control is established;Based on the cost model of building load collaborative control, actual building load collaborative control problem is converted into deep reinforcement learning problem;Iterative solution deep reinforcement learning problem, obtain final strategy, control building in distributed generator and each air conditioner.The state matrix and reward function of the deep reinforcement learning algorithm are improved in the application, and a set of efficient, reliable, privacy-friendly building load collaborative control method is constructed, which provides technical support for the practical application of V2B mode, and helps the efficient consumption of renewable energy and the low-carbon transformation of energy system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of building load management and relates to building load collaborative optimization and control technology, specifically to a building load collaborative control method that considers the needs of V2B mode users. Background Technology

[0002] With the rapid increase in the proportion of renewable energy power generation, its characteristics of strong volatility, high randomness, and weak regulation flexibility have brought unprecedented challenges to the safe and stable operation of the power grid. As an important component of the load side, building electricity load has high flexibility, and coordinated optimization and control of building load can, to some extent, reduce the impact of the uncertainty of renewable energy output on the optimized operation of the system.

[0003] Building upon the framework of coordinated building load regulation, and leveraging the high energy storage capacity of electric vehicles (EVs), a new perspective on energy optimization is provided through a V2B (Vehicle-to-Building) model that enables flexible bidirectional charging and discharging of EVs. In this model, EVs not only function as part of the electrical load but also feed their battery power back into the building's power grid via bidirectional charging and discharging technology. When an EV's battery is charging, charging current flows into the battery; during peak electricity consumption periods, the battery can feed its stored energy back into the building, participating in load regulation and thus achieving a dynamic balance between power supply and demand.

[0004] However, in the V2B model, both building load demand and electric vehicle charging / discharging behavior exhibit high dynamism and stochasticity. Their nonlinear and time-varying characteristics make traditional control methods difficult to handle. Traditional methods typically rely on accurate mathematical modeling and static optimization strategies, which show significant limitations when facing the complex dynamic characteristics of building load and electric vehicle charging / discharging behavior. Furthermore, uncertain user behavior, electricity price fluctuations, and the intermittent nature of renewable energy generation further hinder the real-time performance and flexibility of traditional methods. Deep Reinforcement Learning (DRL), through continuous interaction with the environment, can autonomously learn optimal strategies and dynamically adapt to changing operating environments without relying on precise models. Its powerful online learning capabilities and global optimization characteristics provide a new perspective for solving complex control problems in the V2B model, and can more effectively address the uncertainties and dynamic challenges in building load collaborative optimization.

[0005] Existing research using deep reinforcement learning for collaborative optimization and control of building loads in a V2B model often aims for optimal system operation and incorporates energy storage lifetime loss as part of the loss function. Furthermore, existing deep reinforcement learning-based load collaborative optimization and control methods often use user privacy information such as photovoltaic output, load, air conditioning on / off times, and electric vehicle arrival and departure times as known state parameter inputs.

[0006] Frequent charging and discharging accelerates the performance degradation of electric vehicle batteries, affecting their lifespan. This not only increases battery replacement costs for car owners but may also reduce their acceptance of the V2B model. However, current technologies rarely consider developing appropriate control strategies to ensure the rational use of energy storage devices and ultimately extend their lifespan. Furthermore, in the V2B model, the coordinated control of electric vehicles and loads such as building air conditioning requires the real-time collection and sharing of large amounts of data. Without privacy protection mechanisms, this data could be misused or leaked. However, current research has failed to fully leverage the advantages of deep reinforcement learning algorithms while employing reasonable methods to protect privacy information, further reducing user acceptance of the V2B model. Summary of the Invention

[0007] Purpose of the invention: To address the issues of battery life management and privacy protection in the collaborative optimization and control of building loads under the V2B model, this invention provides a building load collaborative control method that considers the needs of V2B users. It improves the state matrix and reward function of the deep reinforcement learning algorithm, and constructs an efficient, reliable, and privacy-friendly building load collaborative control method. This provides technical support for the practical application of the V2B model and helps to efficiently absorb renewable energy and promote the low-carbon transformation of the energy system.

[0008] Technical Solution: To achieve the above objectives, this invention provides a building load collaborative control method that considers the needs of V2B mode users, comprising the following steps:

[0009] S1: In the coordinated control of building loads, stochastic equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power models of building photovoltaics and rigid loads are obtained.

[0010] S2: Construct a mathematical model for a distributed generator;

[0011] S3: Construct a mathematical model for electric vehicle battery energy storage;

[0012] S4: Construct a mathematical model of air conditioning load;

[0013] S5: Based on the model constructed in steps S1 to S4, establish the power balance constraint equations;

[0014] S6: Based on the power balance constraint equation, establish a cost model for coordinated control of building load. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of loss of user satisfaction in the building due to the failure of electric vehicle energy storage to meet usage needs, and the cost of loss of user satisfaction in the building due to unsuitable indoor temperature.

[0015] S7: A cost model based on building load coordination and control transforms the actual building load coordination and control problem into a deep reinforcement learning problem;

[0016] S8: Iteratively solve the deep reinforcement learning problem to obtain the final policy π. Based on the final policy π, select the corresponding optimal action a(t) to control the distributed generators and air conditioners in the building, so as to realize the coordinated control of building load considering the needs of V2B mode users.

[0017] Furthermore, the creation of low-carbon buildings in step S1, utilizing clean energy primarily based on photovoltaic power generation, is a crucial measure for its green transformation. Building photovoltaics typically exist in a distributed form, generating and consuming electricity locally, reducing power transmission losses and improving energy efficiency. Within buildings, there are some rigid loads that are unsuitable for regulation, such as basic lighting loads, household electricity (kitchen refrigerators, microwave ovens, etc.), and office equipment (computers, printers, etc.). In the coordinated control of building loads, stochastic equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power model of building photovoltaics and rigid loads can be obtained as follows:

[0018]

[0019] P L (t)=μ L P L (t)dt+σ L dω L (t) (2)

[0020] Where t is the time of the system, P PV (t), P L (t) represent the output power of the building's photovoltaic system and rigid load at time t, respectively; P PVT (t) represents the theoretical solar irradiance received by the photovoltaic system at time t, without considering weather conditions. This represents the overall trend in photovoltaic output power; ω PV (t) and ω L (t) is standard Brownian motion; μ PV μ L These are the drift coefficients for building photovoltaic systems and rigid loads, respectively, σ PV σ L These are the diffusion coefficients for building photovoltaics and rigid loads, respectively.

[0021] Further, in step S1, μ PV μ L σ PV σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows:

[0022]

[0023] Where ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, and P(·) is the transition probability function;

[0024] By setting the partial differential derivative to 0, the undetermined system parameter μ is obtained. PV σ PV μ L σ L Specific values:

[0025]

[0026] Furthermore, the mathematical model of the distributed generator in step S2 is as follows:

[0027]

[0028] Among them, P DG (t) represents the generator output power at time t; T DG It is the generator time constant, and the specific value depends on the engine model. This is the generator's maximum output power; u DG (t) represents the generator control signal.

[0029] Furthermore, in step S3, the V2B technology enables bidirectional energy flow between electric vehicles and buildings. Electric vehicles charge when the building's electricity consumption is low and supply power to the building when the electricity consumption is high, which can smooth the load curve and is consistent with the peak shaving and valley filling function of energy storage devices.

[0030] The mathematical model for electric vehicle battery energy storage is as follows:

[0031]

[0032] in, P represents the battery charging / discharging state of the i-th electric vehicle at time t; i BES Q(t) is the battery output power of the i-th electric vehicle at time t; s,i η is the capacity coefficient of the i-th electric vehicle; in,i η out,i These are the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle, respectively.

[0033] Furthermore, in step S4, based on usage and characteristics, building load includes lighting load, air conditioning and cooling load, and household and office equipment load, among which air conditioning is usually the largest single adjustable flexible load in a building, especially in commercial and residential buildings, where its power consumption can account for 30%-50% of the total energy consumption. In addition, air conditioning energy consumption is affected by outdoor temperature, indoor comfort requirements, and user habits, exhibiting strong fluctuations and flexibility. Therefore, the air conditioning load is modeled separately, and the mathematical model for the air conditioning load is:

[0034]

[0035] Changes in indoor temperature are related to outdoor temperature and air conditioner output power, as shown in the following mathematical model:

[0036]

[0037] in, It is the indoor temperature of the j-th room at time t; and It is a coefficient determined by the characteristics of the corresponding room and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

[0038] Furthermore, the power balance constraint equation in step S5 is:

[0039]

[0040] Where, N EV and N AC These represent the number of electric vehicles and air conditioners within the building setting.

[0041] Furthermore, the mathematical expression for the cost of the distributed generator in step S6 is:

[0042]

[0043] Among them, C DG (t) represents the generator cost at time t; and These are the generator cost coefficients;

[0044] By introducing the PLET model, a mathematical expression for the depreciation cost of electric vehicle energy storage is considered. This model represents the energy storage lifetime depreciation as LOH. The total lifespan of energy storage devices is expressed as Where n is the total number of charge-discharge cycles considered by the model within the time range; d i (t) is the depth of discharge in the i-th charge-discharge cycle. k p It is the Peukert lifetime constant, obtained through parameter identification, and is typically in the range of 1.1 to 1.3; for a specific d, It can be considered a constant. Therefore, by limiting ΔC... PLET This can achieve the effect of extending the lifespan of energy storage.

[0045] Therefore, the energy storage lifetime loss cost of the i-th electric vehicle at time t is expressed as c. PLET,i (t):

[0046]

[0047] The cost of lost user satisfaction in the building due to the inability of electric vehicle energy storage to meet usage demand is as follows:

[0048]

[0049] in, The cost of the satisfaction loss of the i-th electric vehicle in the building at time t; and These are the charging cost coefficients for the i-th electric vehicle within the building; This is the difference between the current battery level of the electric vehicle and the user's desired battery level.

[0050] Consider the cost of lost user satisfaction due to unsuitable indoor temperature:

[0051]

[0052] in, The cost of customer satisfaction loss at time t for the j-th air conditioner in the building; T min,j It is the minimum temperature set by the user in the room where the j-th air conditioner is located in the building, T max,j It is the maximum temperature limit set by the user in the room where the j-th air conditioner is located in the building; and The comfort loss weighting coefficient for the j-th air conditioner in the building;

[0053] The cost model for coordinated building load control includes the total cost of system operation and user satisfaction loss. The sum of the costs of distributed generators, energy storage lifespan depreciation, user satisfaction loss due to electric vehicle energy storage not meeting usage demand, and user satisfaction loss due to unsuitable indoor temperatures is the total cost of system operation and user satisfaction loss.

[0054]

[0055] Among them, C ALL (t) represents the total cost.

[0056] Furthermore, step S7 specifically includes:

[0057] The practical problem of coordinated building load control is transformed into a deep reinforcement learning problem. The components are defined as follows: state space S, action space A, reward function r, and agent policy π. In the deep reinforcement learning algorithm, the agent, defined by the algorithm, interacts with the environment, learns the optimal policy π, and optimizes decisions to achieve a specific goal. The agent's main function is to select an appropriate action a(t) based on the current state S(t) and continuously adjust its decision-making process through environmental feedback (reward or punishment) to maximize the expected cumulative reward R(t).

[0058] The state space S(t) contains all the information of the building system at time t, including the photovoltaic output P. PV (t), rigid load P L (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle arrival time Electric vehicle departure time Indoor temperature Air conditioning start time Air conditioner off time It can be represented as:

[0059]

[0060] However, some status information can leak users' personal privacy information. For example, the arrival and departure times of electric vehicles can reflect users' travel habits, the on / off times of air conditioners can indicate whether anyone is controlling them in the building, and the output and load of solar panels can reveal the size of the building and the approximate number of users. For privacy reasons, some information needs to be protected. Therefore, electric vehicle status information is introduced. and air conditioning status in 0 indicates that the air conditioner cannot participate in temperature control (the electric vehicle is not at a charging station or the owner does not wish to participate in temperature control, or the air conditioner does not participate in indoor temperature adjustment); 1 indicates that the air conditioner can participate in temperature control (the electric vehicle participates in temperature control, or the air conditioner participates in indoor temperature adjustment). The power difference P is used. pure (t) represents the difference between photovoltaic output and load, P pure (t)=P PV (t)-P L (t) is used to protect photovoltaic output and load data. Additionally, indoor temperature data is hidden to protect users' temperature preferences.

[0061] Therefore, the state space is modified as follows:

[0062] At this point, the state S′(t) contains all the non-privacy-protected information of the building system at time t, including the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status By defining an appropriate state space, the agent can better perceive changes in the building environment and respond accordingly.

[0063] Controllable devices include distributed generators and building air conditioning; therefore, the motion space is represented as follows:

[0064]

[0065] Among them, u DG (t) is the control signal of the distributed generator. It is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signals of the load aggregator for the distributed generators and each air conditioner at time t;

[0066] The proposed building load coordination and control scheme aims to minimize the overall system operating cost while ensuring user experience. In this sense, the reward at time t, which is essentially the reward for higher system operation and user satisfaction losses, is lower. Therefore, the reward at time t is:

[0067]

[0068] At each time t, the agent obtains the building system state S′(t), and from the state S′(t) knows the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information are used to select action a(t) from action space A according to policy π, and to obtain control signals for distributed generators and building air conditioning from a(t), where π: Ω→A is the mapping from state S′(t) to action a(t), which describes the probability that the agent selects action a(t) given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and this process continues until the terminal time T. The discounted cumulative reward starting from time t is defined as:

[0069]

[0070] Where γ is the discount factor, γ∈(0,1).

[0071] The goal of deep reinforcement learning algorithms is to learn a policy π that maximizes the initial reward J[π] = E[R(0)]maxJ[π]. This is because a larger reward r(t) at each step represents a lower system cost. Therefore, when the initial reward is maximized, the total cost of the entire building system operation is minimized.

[0072] Furthermore, the method for obtaining the final strategy π in step S8 is as follows:

[0073] The policy π of all agents is trained using deep reinforcement learning to maximize the objective function value J[π]. The policy π describes the probability of an agent choosing action a(t) in a given state S′(t). By training the policy π, the agent chooses the action that maximizes the objective function in each state according to the policy π, thereby minimizing the operating cost of the entire system while ensuring user experience.

[0074] To better represent the policy, the value function V is defined under the policy π and the state value S′(t). π (S′(t))=E[R(t)|S′(t)]; where E[●] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under policy π; the value function can be used to measure the quality of S′(t), if V π A high (S′(t)) indicates that starting from this state, the agent can obtain a higher reward by acting according to policy π. In building load regulation, the value function V π (S′(t)) can be used to evaluate whether the current building state is beneficial to the objective. High-value states guide the agent to prioritize maintaining or transitioning to more favorable building states (i.e., minimizing the overall system operating cost while ensuring user experience). Based on the value function, for each time t, the advantage function is calculated. The parameters θ and θ' used to update the policy function and value function v ;θ and θ v The policy function π(·|·;θ) and the value function V(·;θ) are the policy function and the value function V(·;θ). v An important parameter in π is θ, which is the policy parameter. By optimizing θ, policy π can learn how to select the optimal action in different buildings and under different conditions. v These are the parameters of the value function, used to evaluate the overall performance in the current state (whether it reduces the operating costs of the building system and whether it guarantees the user experience); parameters θ and θ' v Update using the following methods:

[0075]

[0076] in, α represents the gradient of the policy function with respect to the parameter θ, used to adjust the probability of choosing an action; α and β are the parameters θ and θ' respectively. v The learning rate determines the step size for each update;

[0077] By iteratively updating the parameter θ v This improves the accuracy of state and action value estimation; based on value function feedback, it optimizes the policy parameter θ to increase the selection probability of high-dominance actions; it repeats the above process until the output distribution of the policy function π(·|·θ) remains almost unchanged in multiple consecutive iterations, indicating that the policy has stabilized and the value function V(·θ) has reached its optimal level. v The error of the load aggregator approaches zero, indicating that the value assessment is sufficiently accurate. During the iteration process, the strategy π(·|·θ) gradually tends to select the action with the highest return in all states. As the iteration is completed, the load aggregator obtains the final strategy π. Based on the obtained strategy π, the load aggregator can, at any time, use the observed building system state S′(t) (the difference between photovoltaic output and load P) to determine the optimal action. L (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status The corresponding optimal action a(t) is selected to control the distributed generators and air conditioners in the building, thereby realizing the coordinated control of building load considering the needs of V2B mode users.

[0078] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0079] 1. This invention optimizes the charging and discharging strategy of electric vehicle batteries by improving the reward function in deep reinforcement learning algorithms, thereby rationally controlling the charging and discharging frequency and depth, and reducing the performance degradation caused by overuse. This method effectively extends the service life of energy storage devices, reduces the economic costs associated with battery replacement, and thus increases user acceptance of the V2B model.

[0080] 2. This invention designs a load control scheme that can protect user privacy. By improving the state matrix of the deep reinforcement learning algorithm, it effectively avoids the algorithm's excessive reliance on sensitive user data (arrival and departure times of electric vehicles, photovoltaic power output, building load, and air conditioning on / off times in the building), thereby reducing the risk of privacy leakage and enhancing users' trust in this V2B technology. Attached Figure Description

[0081] Figure 1 A schematic diagram illustrating a building load optimization and control implementation scenario under the V2B model;

[0082] Figure 2 This is a flowchart illustrating the method of the present invention. Detailed Implementation

[0083] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0084] like Figure 1 As shown, this invention considers a building load optimization and control scheme based on deep reinforcement learning. In this scenario, a central load aggregator interacts with all distributed buildings and resources through a deep reinforcement learning agent. This agent receives real-time data from buildings (such as load demand, distributed energy generation, and electric vehicle charging / discharging status) and external signals (such as weather data), forming the system status input. Each building unit has local energy dispatch capabilities (such as self-consumption of solar energy, air conditioning load adjustment, and electric vehicle charging / discharging management). Through deep reinforcement learning, the central control system can collaboratively optimize the load distribution of multiple buildings, dynamically allocate resources, and achieve energy balance and maximize utilization efficiency. In the figure, electric vehicles, as mobile energy storage units, are connected to the building system via a V2B model. Deep reinforcement learning can predict the usage behavior and battery status of electric vehicles, making charging / discharging decisions at appropriate times, thereby mitigating building load fluctuations and supporting grid operation. The deep reinforcement learning algorithm continuously updates the strategy function through continuous environmental interaction (such as real-time building load data), enabling the control model to adapt to the dynamic changes in building load and grid conditions.

[0085] Based on the above implementation scenarios, this invention provides a building load collaborative control method that considers the needs of V2B mode users, such as... Figure 2 As shown, it includes the following steps:

[0086] S1: In the coordinated control of building loads, stochastic equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power models of building photovoltaics and rigid loads are obtained.

[0087] Building low-carbon buildings and using clean energy, primarily photovoltaic power generation, is a crucial step in their green transformation. Building photovoltaic systems typically exist in a distributed manner, generating and consuming electricity locally, reducing transmission losses and improving energy efficiency. Within buildings, there are some rigid loads that are unsuitable for regulation, such as basic lighting loads, household electricity (kitchen refrigerators, microwave ovens, etc.), and office equipment (computers, printers, etc.). In building load coordinated regulation, stochastic equations are used to simulate the power changes of building photovoltaic systems and rigid loads. After a series of feedback approximations, the power model of building photovoltaic systems and rigid loads can be obtained as follows:

[0088]

[0089] P L (t)=μ L P L (t)dt+σ L dω L (t) (2)

[0090] Where t is the time of the system, P PV (t), P L (t) represent the output power of the building's photovoltaic system and rigid load at time t, respectively; P PVT (t) represents the theoretical solar irradiance received by the photovoltaic system at time t, without considering weather conditions. This represents the overall trend in photovoltaic output power; ω PV (t) and ω L (t) is standard Brownian motion; μ PV μ L These are the drift coefficients for building photovoltaic systems and rigid loads, respectively, σ PV σ L These are the diffusion coefficients for building photovoltaic systems and rigid loads, respectively.

[0091] μ PV μ L σ PV σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows:

[0092]

[0093] Where ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, and P(·) is the transition probability function;

[0094] By setting the partial differential derivative to 0, the undetermined system parameter μ is obtained. PV σ PV μ L σ L Specific values:

[0095]

[0096] S2: Constructing a mathematical model for a distributed generator:

[0097]

[0098] Among them, P DG (t) represents the generator output power at time t; T DG It is the generator time constant, and the specific value depends on the engine model. This is the generator's maximum output power; u DG (t) represents the generator control signal.

[0099] S3: Constructing a mathematical model for electric vehicle battery energy storage:

[0100] V2B technology enables bidirectional energy flow between electric vehicles and buildings. Electric vehicles charge when the building's electricity consumption is low and supply power to the building when the electricity consumption is high, which can smooth the load curve and is consistent with the peak shaving and valley filling function of energy storage devices.

[0101] The mathematical model for electric vehicle battery energy storage is as follows:

[0102]

[0103] in, P represents the battery charging / discharging state of the i-th electric vehicle at time t; i BES Q(t) is the battery output power of the i-th electric vehicle at time t; s,i η is the capacity coefficient of the i-th electric vehicle; in,i η out,i These are the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle, respectively.

[0104] S4: Constructing a mathematical model for air conditioning load:

[0105] Based on their purpose and characteristics, building loads include lighting loads, air conditioning and cooling loads, and household and office equipment loads. Air conditioning is typically the largest single adjustable flexible load in a building, especially in commercial and residential buildings, where it can account for 30%-50% of total energy consumption. Furthermore, air conditioning energy consumption is significantly affected by outdoor temperature, indoor comfort requirements, and user habits, exhibiting considerable fluctuation and flexibility. Therefore, a separate mathematical model for air conditioning load is developed:

[0106]

[0107] Changes in indoor temperature are related to outdoor temperature and air conditioner output power, as shown in the following mathematical model:

[0108]

[0109] in, It is the indoor temperature of the j-th room at time t; and It is a coefficient determined by the characteristics of the corresponding room and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

[0110] S5: Based on the model constructed in steps S1 to S4, establish the power balance constraint equations:

[0111]

[0112] Where, N EV and N AC These represent the number of electric vehicles and air conditioners within the building setting.

[0113] S6: Based on the power balance constraint equation, establish a cost model for coordinated control of building load. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of loss of user satisfaction in the building due to the failure of electric vehicle energy storage to meet usage needs, and the cost of loss of user satisfaction in the building due to unsuitable indoor temperature.

[0114] The mathematical expression for the cost of a distributed generator is:

[0115]

[0116] Among them, C DG (t) represents the generator cost at time t; and These are the generator cost coefficients;

[0117] By introducing the PLET model, a mathematical expression for the depreciation cost of electric vehicle energy storage is considered. This model represents the energy storage lifetime depreciation as LOH. The total lifespan of energy storage devices is expressed as Where n is the total number of charge-discharge cycles considered by the model within the time range; d i (t) is the depth of discharge in the i-th charge-discharge cycle. k p It is the Peukert lifetime constant, obtained through parameter identification, and is typically in the range of 1.1 to 1.3; for a specific d, It can be considered a constant. Therefore, by limiting ΔC... PLET This can achieve the effect of extending the lifespan of energy storage.

[0118] Therefore, the energy storage lifetime loss cost of the i-th electric vehicle at time t is expressed as c. PLET,i (t):

[0119]

[0120] The cost of lost user satisfaction in the building due to the inability of electric vehicle energy storage to meet usage demand is as follows:

[0121]

[0122] in, The cost of the satisfaction loss of the i-th electric vehicle in the building at time t; and These are the charging cost coefficients for the i-th electric vehicle within the building; This is the difference between the current battery level of the electric vehicle and the user's desired battery level.

[0123] Consider the cost of lost user satisfaction due to unsuitable indoor temperature:

[0124]

[0125] in, The cost of customer satisfaction loss at time t for the j-th air conditioner in the building; T min,j It is the minimum temperature set by the user in the room where the j-th air conditioner is located in the building, T max,j It is the maximum temperature limit set by the user in the room where the j-th air conditioner is located in the building; and The comfort loss weighting coefficient for the j-th air conditioner in the building;

[0126] The cost model for coordinated building load control includes the total cost of system operation and user satisfaction loss. The sum of the costs of distributed generators, energy storage lifespan depreciation, user satisfaction loss due to electric vehicle energy storage not meeting usage demand, and user satisfaction loss due to unsuitable indoor temperatures is the total cost of system operation and user satisfaction loss.

[0127]

[0128] Among them, C ALL (t) represents the total cost.

[0129] S7: A cost model based on building load coordination and control transforms the actual building load coordination and control problem into a deep reinforcement learning problem;

[0130] The practical problem of coordinated building load control is transformed into a deep reinforcement learning problem. The components are defined as follows: state space S, action space A, reward function r, and agent policy π. In the deep reinforcement learning algorithm, the agent, defined by the algorithm, interacts with the environment, learns the optimal policy π, and optimizes decisions to achieve a specific goal. The agent's main function is to select an appropriate action a(t) based on the current state S(t) and continuously adjust its decision-making process through environmental feedback (reward or punishment) to maximize the expected cumulative reward R(t).

[0131] The state space S(t) contains all the information of the building system at time t, including the photovoltaic output P. PV (t), rigid load P L (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle arrival time Electric vehicle departure time Indoor temperature Air conditioning start time Air conditioner off time It can be represented as:

[0132]

[0133] However, some status information can leak users' personal privacy information. For example, the arrival and departure times of electric vehicles can reflect users' travel habits, the on / off times of air conditioners can indicate whether anyone is controlling them in the building, and the output and load of solar panels can reveal the size of the building and the approximate number of users. For privacy reasons, some information needs to be protected. Therefore, electric vehicle status information is introduced. and air conditioning status in 0 indicates that the air conditioner cannot participate in temperature control (the electric vehicle is not at a charging station or the owner does not wish to participate in temperature control, or the air conditioner does not participate in indoor temperature adjustment); 1 indicates that the air conditioner can participate in temperature control (the electric vehicle participates in temperature control, or the air conditioner participates in indoor temperature adjustment). The power difference P is used. pure(t) represents the difference between photovoltaic output and load, P pure (t)=P PV (t)-P L (t) is used to protect photovoltaic output and load data. Additionally, indoor temperature data is hidden to protect users' temperature preferences.

[0134] Therefore, the state space is modified as follows:

[0135] At this point, the state S′(t) contains all the non-privacy-protected information of the building system at time t, including the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status By defining an appropriate state space, the agent can better perceive changes in the building environment and respond accordingly.

[0136] Controllable devices include distributed generators and building air conditioning; therefore, the motion space is represented as follows:

[0137]

[0138] Among them, u DG (t) is the control signal of the distributed generator. It is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signals of the load aggregator for the distributed generators and each air conditioner at time t;

[0139] The proposed building load coordination and control scheme aims to minimize the overall system operating cost while ensuring user experience. In this sense, the reward at time t, which is essentially the reward for higher system operation and user satisfaction losses, is lower. Therefore, the reward at time t is:

[0140]

[0141] At each time t, the agent obtains the building system state S′(t), and from the state S′(t) knows the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information are used to select action a(t) from action space A according to policy π, and to obtain control signals for distributed generators and building air conditioning from a(t), where π: Ω→A is the mapping from state S′(t) to action a(t), which describes the probability that the agent selects action a(t) given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and this process continues until the terminal time T. The discounted cumulative reward starting from time t is defined as:

[0142]

[0143] Where γ is the discount factor, γ∈(0,1).

[0144] The goal of deep reinforcement learning algorithms is to learn a policy π that maximizes the initial reward J[π] = E[R(0)]maxJ[π]. This is because a larger reward r(t) at each step represents a lower system cost. Therefore, when the initial reward is maximized, the total cost of the entire building system operation is minimized.

[0145] S8: Iteratively solve the deep reinforcement learning problem to obtain the final policy π. Based on the final policy π, select the corresponding optimal action a(t) to control the distributed generators and air conditioners in the building, realizing coordinated control of building load considering the needs of V2B users. Specifically, this includes:

[0146] The final strategy π is obtained as follows:

[0147] The policy π of all agents is trained using deep reinforcement learning methods to maximize the objective function value. Policy π describes the probability that an agent will choose action a(t) given a state S′(t). By training policy π, the agent will choose the action that maximizes the objective function in each state, thereby minimizing the operating cost of the entire system while ensuring user experience.

[0148] To better represent the policy, the value function V is defined under the policy π and the state value S′(t). π (S′(t))=E[R(t)|S′(t)]; where E[·] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under policy π; the value function can be used to measure the quality of S′(t), if V π A high (S′(t)) indicates that starting from this state, the agent can obtain a higher reward by acting according to policy π. In building load regulation, the value function V π(S′(t)) can be used to evaluate whether the current building state is beneficial to the objective. High-value states guide the agent to prioritize maintaining or transitioning to more favorable building states (i.e., minimizing the overall system operating cost while ensuring user experience). Based on the value function, for each time t, the advantage function is calculated. The parameters θ and θ' used to update the policy function and value function v ;θ and θ v The policy function π(·|·;θ) and the value function V(·;θ) are the policy function and the value function V(·;θ). v An important parameter in π is θ, which is the policy parameter. By optimizing θ, policy π can learn how to select the optimal action in different buildings and under different conditions. v These are the parameters of the value function, used to evaluate the overall performance in the current state (whether it reduces the operating costs of the building system and whether it guarantees the user experience); parameters θ and θ' v Update using the following methods:

[0149]

[0150] in, α represents the gradient of the policy function with respect to the parameter θ, used to adjust the probability of choosing an action; α and β are the parameters θ and θ' respectively. v The learning rate determines the step size for each update;

[0151] By iteratively updating the parameter θ v This improves the accuracy of state and action value estimation; based on value function feedback, it optimizes the policy parameter θ to increase the selection probability of high-dominance actions; it repeats the above process until the output distribution of the policy function π(·|·θ) remains almost unchanged in multiple consecutive iterations, indicating that the policy has stabilized and the value function V(·; θ) is stable. v The error of the load aggregator approaches zero, indicating that the value assessment is sufficiently accurate. During the iteration process, the strategy π(·|·θ) gradually tends to select the action with the highest return in all states. As the iteration is completed, the load aggregator obtains the final strategy π. Based on the obtained strategy π, the load aggregator can, at any time, use the observed building system state S′(t) (the difference between photovoltaic output and load P) to determine the optimal action. L (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status The corresponding optimal action a(t) is selected to control the distributed generators and air conditioners in the building, thereby realizing the coordinated control of building load considering the needs of V2B mode users.

[0152] Based on the above scheme, in order to verify the actual effect of the method of the present invention, this embodiment compares the control method of the present invention with existing control methods in an experiment, as follows:

[0153] The experiment yielded comparative data on battery life degradation, as shown in Table 1.

[0154] Table 1

[0155]

[0156] As can be seen from Table 1, the battery life loss under the control method of the present invention is better than that of the other three methods, and the method of the present invention can extend the battery life.

Claims

1. A building load collaborative control method considering the needs of V2B mode users, characterized in that, Includes the following steps: S1: In the coordinated control of building loads, stochastic equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power models of building photovoltaics and rigid loads are obtained. S2: Construct a mathematical model for a distributed generator; S3: Construct a mathematical model for electric vehicle battery energy storage; S4: Construct a mathematical model of air conditioning load; S5: Based on the model constructed in steps S1 to S4, establish the power balance constraint equations; S6: Based on the power balance constraint equation, establish a cost model for coordinated control of building load. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of loss of user satisfaction in the building due to the failure of electric vehicle energy storage to meet usage needs, and the cost of loss of user satisfaction in the building due to unsuitable indoor temperature. S7: A cost model based on building load coordination and control transforms the actual building load coordination and control problem into a deep reinforcement learning problem; S8: Iteratively solve the deep reinforcement learning problem to obtain the final policy π. Based on the final policy π, select the corresponding optimal action a(t) to control the distributed generators and air conditioners in the building, so as to realize the coordinated control of building load considering the needs of V2B mode users.

2. The building load collaborative control method considering V2B mode user demand according to claim 1, characterized in that, The power models for building photovoltaics and rigid loads in step S1 are as follows: P L (t)=μ L P L (t)dt+σ L dω L (t) (2) Where t is the time of the system, P PV (t), P L (t) represent the output power of the building's photovoltaic system and rigid load at time t, respectively; P PVT (t) represents the theoretical solar irradiance received by the photovoltaic system at time t, without considering weather conditions. This represents the overall trend in photovoltaic output power; ω PV (t) and ω L (t) is standard Brownian motion; μ PV μ L These are the drift coefficients for building photovoltaic systems and rigid loads, respectively, σ PV σ L These are the diffusion coefficients for building photovoltaics and rigid loads, respectively.

3. A building load collaborative control method considering V2B mode user needs according to claim 2, characterized in that, In step S1, μ PV μ L σ PV σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows: Where ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, and P(·) is the transition probability function; By setting the partial differential derivative to 0, the undetermined system parameter μ is obtained. PV σ PV μ L σ L Specific values:

4. A building load collaborative control method considering V2B mode user needs according to claim 1, characterized in that, The mathematical model of the distributed generator in step S2 is as follows: Among them, P DG (t) represents the generator output power at time t; T DG It is the generator time constant; This is the generator's maximum output power; u DG (t) represents the generator control signal.

5. A building load collaborative control method considering V2B mode user demand according to claim 1, characterized in that, The mathematical model for electric vehicle battery energy storage in step S3 is as follows: in, P represents the battery charging / discharging state of the i-th electric vehicle at time t; i BES Q(t) is the battery output power of the i-th electric vehicle at time t; s,i η is the capacity coefficient of the i-th electric vehicle; in,i η out,i These are the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle, respectively.

6. A building load collaborative control method considering V2B mode user needs according to claim 1, characterized in that, The mathematical model for air conditioning load in step S4 is as follows: Changes in indoor temperature are related to outdoor temperature and air conditioner output power, as shown in the following mathematical model: in, It is the indoor temperature of the j-th room at time t; and It is a coefficient determined by the characteristics of the corresponding room and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

7. A building load collaborative control method considering V2B mode user demand according to claim 1, characterized in that, The power balance constraint equation in step S5 is: Where, N EV and N AC These represent the number of electric vehicles and air conditioners within the building scene, respectively; P i BES (t) represents the battery output power of the i-th electric vehicle at time t; P L (t), P PV (t) represents the output power of the building's photovoltaic system and rigid load at time t, respectively; P DG (t) represents the generator output power at time t.

8. A building load collaborative control method considering V2B mode user demand according to claim 1, characterized in that, The mathematical expression for the cost of the distributed generator in step S6 is: Among them, C DG (t) represents the generator cost at time t; and These are the generator cost coefficients; P DG (t) represents the generator output power at time t. By introducing the PLET model, the energy storage lifetime depreciation cost of the i-th electric vehicle at time t is expressed as c. PLET,i (t): Among them, P i BES (t) represents the battery output power of the i-th electric vehicle at time t; Q s,i η is the capacity coefficient of the i-th electric vehicle; out,i Let be the discharge efficiency coefficient of the i-th electric vehicle; The cost of lost user satisfaction in the building due to the inability of electric vehicle energy storage to meet usage demand is as follows: in, The cost of the satisfaction loss of the i-th electric vehicle in the building at time t; and These are the charging cost coefficients for the i-th electric vehicle within the building; This is the difference between the current battery level of the electric vehicle and the user's desired battery level. Consider the cost of lost user satisfaction due to unsuitable indoor temperature: in, The cost of customer satisfaction loss at time t for the j-th air conditioner in the building; T min,j It is the minimum temperature set by the user in the room where the j-th air conditioner is located in the building, T max,j It is the maximum temperature limit set by the user in the room where the j-th air conditioner is located in the building; and The comfort loss weighting coefficient for the j-th air conditioner in the building; The cost model for coordinated building load control includes the total cost of system operation and user satisfaction loss: Among them, C ALL (t) represents the total cost.

9. A building load collaborative control method considering V2B mode user demand according to claim 8, characterized in that, Step S7 specifically includes: Define the components, including the state space S, action space A, reward function r, and agent policy π. In the deep reinforcement learning algorithm, the agent, defined by the algorithm, interacts with the environment, learns the optimal policy π, and optimizes decisions to achieve a specific goal within the deep reinforcement learning framework. The agent selects an appropriate action a(t) based on the current state S(t) and continuously adjusts its decision-making process through environmental feedback to maximize the expected cumulative reward R(t). The state space S(t) is represented as: Modify the state space as follows: At this point, the state S′(t) contains all the non-privacy-protected information of the building system at time t, including the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status The action space is represented as: Among them, u DG (t) is the control signal of the distributed generator. It is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signals of the load aggregator for the distributed generators and each air conditioner at time t; The reward at time t is: At each time t, the agent obtains the building system state S′(t), and from the state S′(t) knows the difference P between photovoltaic output and load. pure (t), Output power of distributed generator P DG (t), external temperature T OUT (t) Electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information are used to select action a(t) from action space A according to policy π, and to obtain control signals for distributed generators and building air conditioning from a(t), where π: Ω→A is the mapping from state S′(t) to action a(t), which describes the probability that the agent selects action a(t) given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and this process continues until the terminal time T. The discounted cumulative reward starting from time t is defined as: Where γ is the discount factor, γ∈(0,1).

10. A building load collaborative control method considering V2B mode user demand according to claim 8, characterized in that, The method for obtaining the final strategy π in step S8 is as follows: The policy π of all agents is trained using deep reinforcement learning to maximize the objective function value J[π]. The policy π describes the probability of an agent choosing action a(t) in a given state S′(t). By training the policy π, the agent chooses the action that maximizes the objective function in each state according to the policy π, thereby minimizing the operating cost of the entire system while ensuring user experience. To better represent the policy, the value function V is defined under the policy π and the state value S′(t). π (S′(t))=E[R(t)|S′(t)]; where E[●] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under policy π; based on the value function, for each time t, the advantage function is calculated. The parameters θ and θ' used to update the policy function and value function v ; parameters θ and θ v Update using the following methods: in, α represents the gradient of the policy function with respect to the parameter θ, used to adjust the probability of choosing an action; α and β are the parameters θ and θ' respectively. v The learning rate determines the step size for each update; By iteratively updating the parameter θ v This improves the accuracy of state and action value estimation; based on value function feedback, the policy parameter θ is optimized to increase the selection probability of high-advantage actions; as the iteration is completed, the load aggregator obtains the final policy π.

Citation Information

Patent Citations

  • Power distribution network layered and partitioned load shedding coordination control method based on event trigger mechanism

    CN114844050A

  • Intelligent building electricity demand side optimization scheduling strategy based on particle swarm optimization

    CN116579571A