Building load coordinated regulation and control method considering V2B mode user demand

By improving the state matrix and reward function of the deep reinforcement learning algorithm, the charging and discharging strategies of electric vehicle batteries are optimized, and the privacy protection solution is designed, the battery life management and privacy protection problems in the coordinated control of building loads in V2B mode are solved, and the efficient, reliable, and privacy-friendly load regulation effect is achieved.

CN120016548AActive Publication Date: 2025-05-16HOHAI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510081810.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In V2B mode, building load demand and electric vehicle charging and discharge behavior are highly dynamic and random, which are difficult to cope with, and the existing technology fails to effectively consider battery life management and privacy protection.

Method used

By improving the state matrix and reward function of the deep reinforcement learning algorithm, an efficient, reliable, and privacy-friendly coordinated control method for building loads is constructed, the charging and discharging strategies of electric vehicle batteries are optimized, the loss caused by excessive use is reduced to battery performance, and a load regulation scheme that protects user privacy is designed.

Benefits of technology

It effectively extends the service life of energy storage equipment, reduces the economic costs caused by battery replacement, increases users' acceptance of V2B mode, reduces the risk of privacy leakage, and enhances user trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016548A_ABST
    Figure CN120016548A_ABST
Patent Text Reader

Abstract

The invention discloses a building load coordinated regulation and control method considering V2B mode user demands. The method comprises the following steps: obtaining a power model of building photovoltaic and rigid loads; constructing a mathematical model of the distributed generator; constructing an electric vehicle battery energy storage mathematical model; constructing an air conditioner load mathematical model; establishing a power balance constraint equation; based on the power balance constraint equation, establishing a cost model of building load coordinated regulation and control; converting an actual building load coordinated regulation problem into a deep reinforcement learning problem based on a building load coordinated regulation cost model; and iteratively solving the deep reinforcement learning problem to obtain a final strategy, and controlling a distributed generator and each air conditioner in the building. According to the invention, a state matrix and a reward function of a deep reinforcement learning algorithm are improved, a set of efficient, reliable and privacy-friendly building load coordinated regulation and control method is constructed, technical support is provided for practical application of a V2B mode, and efficient consumption of renewable energy and low-carbon transformation of an energy system are assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of building load management, and relates to a building load collaborative optimization and control technology, and specifically to a building load collaborative control method that takes into account V2B mode user needs. Background Art

[0002] With the rapid increase in the proportion of renewable energy generation, its strong volatility, strong randomness and weak regulation flexibility have brought unprecedented challenges to the safe and stable operation of the power grid. As an important part of the load side, building power loads are highly flexible. Coordinated optimization and regulation of building loads can reduce the impact of renewable energy output uncertainty on system optimization operations to a certain extent.

[0003] On the basis of coordinated control of building loads, the high energy storage capacity of electric vehicles is utilized to realize flexible bidirectional charging and discharging of electric vehicles through the V2B (Vehicle-to-Building) mode, providing a new perspective for energy optimization. In this mode, electric vehicles are not only part of the power load, but can also feed back the power in their batteries to the building power grid through bidirectional charging and discharging technology. When the battery of an electric vehicle is charging, the charging current flows into the battery of the electric vehicle; during peak hours of electricity consumption, the battery of the electric vehicle can feed back the stored electric energy to the building, participate in the load regulation of the building, and thus achieve a dynamic balance of power supply and demand.

[0004] However, in the V2B mode, building load demand and electric vehicle charging and discharging behavior are highly dynamic and random, and their nonlinear and time-varying characteristics make it difficult for traditional control methods to cope with them. Traditional methods usually rely on accurate mathematical modeling and static optimization strategies, which show obvious limitations when facing the complex dynamic characteristics of building loads and electric vehicle charging and discharging behaviors. In addition, uncertain user behavior, electricity price fluctuations, and intermittent renewable energy generation make traditional methods insufficient in terms of real-time and flexibility. Deep reinforcement learning (DRL) can autonomously learn optimal strategies and dynamically adapt to changing operating environments without relying on precise models through continuous interaction with the environment. Its powerful online learning ability and global optimization characteristics provide a new perspective for solving complex control problems in the V2B mode, and can more effectively cope with the uncertainty and dynamic challenges in the collaborative optimization of building loads.

[0005] Existing studies using deep reinforcement learning to coordinate and optimize building loads in the V2B mode often take the optimal economic operation of the system as the goal and use the energy storage life loss as part of the loss function. In addition, existing load coordination and optimization control methods based on deep reinforcement learning often use user privacy information such as photovoltaic output, load, air conditioning switching time, and electric vehicle arrival and departure time as known state parameter inputs.

[0006] Frequent charging and discharging will accelerate the performance degradation of electric vehicle batteries and affect their service life. This will not only increase the battery replacement costs of car owners, but may also lead to a decrease in car owners' acceptance of the V2B model. However, existing technologies rarely consider formulating appropriate control strategies so that energy storage equipment can be used reasonably, ultimately achieving the purpose of extending the service life of energy storage equipment. In addition, in the V2B model, the coordinated control of electric vehicles and loads such as air conditioners in buildings requires real-time collection and sharing of a large amount of data. If there is a lack of privacy protection mechanism, this data may be abused or leaked. However, current research has failed to fully utilize the advantages of deep reinforcement learning algorithms while adopting reasonable methods to protect privacy information, which further reduces users' acceptance of the V2B model. Summary of the invention

[0007] Purpose of the invention: To address the issues of battery life management and privacy protection in the collaborative optimization and control of building loads under the V2B mode, a building load collaborative control method that takes into account the user needs of the V2B mode is provided, the state matrix and reward function of the deep reinforcement learning algorithm are improved, and a set of efficient, reliable, and privacy-friendly building load collaborative control methods is constructed to provide technical support for the practical application of the V2B mode, and to help the efficient consumption of renewable energy and the low-carbon transformation of the energy system.

[0008] Technical solution: To achieve the above purpose, the present invention provides a building load coordinated control method considering the needs of V2B mode users, comprising the following steps:

[0009] S1: In the coordinated control of building loads, random equations are used to simulate the power changes of building photovoltaics and rigid loads, and the power models of building photovoltaics and rigid loads are obtained after a series of feedback approximations;

[0010] S2: Construct mathematical model of distributed generators;

[0011] S3: Construct a mathematical model for electric vehicle battery energy storage;

[0012] S4: Construct mathematical model of air conditioning load;

[0013] S5: Establish a power balance constraint equation according to the model constructed in steps S1 to S4;

[0014] S6: Based on the power balance constraint equation, a cost model for coordinated control of building loads is established. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of user satisfaction loss in the building due to insufficient demand for electric vehicle energy storage, and the cost of user satisfaction loss in the building due to inappropriate indoor temperature.

[0015] S7: Based on the cost model of building load collaborative control, the actual building load collaborative control problem is transformed into a deep reinforcement learning problem;

[0016] S8: Iteratively solve the deep reinforcement learning problem to obtain the final strategy π. According to the final strategy π, the corresponding optimal action a(t) is selected to control the distributed generators and air conditioners in the building, so as to realize the coordinated control of building loads considering the user needs of the V2B mode.

[0017] Furthermore, the creation of low-carbon buildings in step S1 and the use of clean energy mainly based on photovoltaic power generation are important measures for its green transformation. Building photovoltaics usually exist in a distributed form, which can generate electricity and consume on-site, reducing power transmission losses and improving energy efficiency. In buildings, there are some rigid loads that are not suitable for regulation, such as basic lighting loads, household electricity (kitchen refrigerators, microwave ovens, etc.) and office equipment (computers, printers, etc.). In the coordinated regulation of building loads, random equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power model of building photovoltaics and rigid loads can be obtained as follows:

[0018]

[0019] P L (t) = μ L P L (t)dt+σ L dω L (t) (2)

[0020] Among them, t is the time when the system is in, P PV (t), P L (t) are the output power of building photovoltaic and rigid load at time t; P PVT (t) is the theoretical solar irradiance received by the photovoltaic at time t without considering weather conditions, is the overall trend of photovoltaic output power; ω PV (t) and ω L (t) is the standard Brownian motion; μ PV , μ L are the drift coefficients of building photovoltaics and rigid loads, σ PV , σ L are the diffusion coefficients of building photovoltaics and rigid loads, respectively.

[0021] Furthermore, in step S1, μ PV , μ L , σ PV , σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows:

[0022]

[0023] Among them, ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, is the transition probability function;

[0024] By setting the partial differential derivative to 0, we can obtain the unknown system parameter μ PV , σ PV , μ L , σ L Specific values:

[0025]

[0026] Furthermore, the mathematical model of the distributed generator in step S2 is:

[0027]

[0028] Among them, P DG (t) is the generator output power at time t; T DG is the generator time constant, the specific value depends on the engine model; is the maximum output power of the generator; u DG (t) is the generator control signal.

[0029] Furthermore, the V2B technology in step S3 realizes a two-way flow of energy between electric vehicles and buildings. Electric vehicles are charged when the building's electricity consumption is low, and power is supplied to the building when the electricity consumption is peak, which can smooth the load curve, consistent with the peak-shaving and valley-filling function of energy storage equipment.

[0030] The mathematical model of electric vehicle battery energy storage is:

[0031]

[0032] in, is the battery charge and discharge status of the i-th electric car at time t; P i BES (t) is the battery output power of the i-th electric vehicle at time t; Q s,i is the capacity factor of the i-th electric vehicle; η in,i , η out,iare the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle respectively.

[0033] Furthermore, in step S4, according to the purpose and characteristics, the building load includes lighting load, air conditioning and refrigeration load, life and office equipment load, etc., among which air conditioning is usually the largest single adjustable flexible load in the building, especially in commercial buildings and residences, and its power consumption accounts for 30%-50% of the total energy consumption. In addition, the energy consumption of air conditioning is affected by outdoor temperature, indoor comfort requirements and user usage habits, showing strong volatility and flexibility. Therefore, the air conditioning load is modeled separately, and the mathematical model of the air conditioning load is:

[0034]

[0035] The change of indoor temperature is related to the outdoor temperature and the output power of the air conditioner. The mathematical model is as follows:

[0036]

[0037] in, is the indoor temperature of the jth room at time t; and is a coefficient determined by the corresponding room characteristics and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

[0038] Furthermore, the power balance constraint equation in step S5 is:

[0039]

[0040] Among them, N EV and N AC are the number of electric vehicles and air conditioners in the building scene, respectively.

[0041] Furthermore, the mathematical expression of the distributed generator cost in step S6 is:

[0042]

[0043] Among them, C DG (t) is the generator cost at time t; and are the generator cost coefficients respectively;

[0044] By introducing the PLET model, the mathematical expression of the energy storage loss cost of electric vehicles is considered. This model expresses the energy storage life loss as LOH, The total life of the energy storage device is expressed as Where n is the total number of charges and discharges within the time range considered by the model; d i(t) is the depth of discharge in the ith charge-discharge cycle, k p is the Peukert life constant, obtained by parameter identification, usually in the range of 1.1 to 1.3; for a specific d, can be considered as a constant. Therefore, by limiting ΔC PLET , which can achieve the effect of extending the energy storage life.

[0045] Therefore, the energy storage life loss cost of the i-th electric vehicle at time t is expressed as c PLET,i (t):

[0046]

[0047] Considering the loss of user satisfaction in the building due to insufficient energy storage demand of electric vehicles, the cost is:

[0048]

[0049] in, is the satisfaction loss cost of the i-th electric car in the building at time t; and are the charging cost coefficients of the i-th electric vehicle in the building; The difference between the current power of the electric vehicle and the power expected by the user;

[0050] Consider the cost of lost building occupant satisfaction due to inappropriate indoor temperature:

[0051]

[0052] in, is the satisfaction loss cost of the jth air conditioner in the building at time t; T min,j is the minimum temperature limit set by the user in the room where the jth air conditioner is located in the building, T max,j is the maximum temperature limit set by the user in the room where the jth air conditioner is located in the building; and is the comfort loss weight coefficient of the jth air conditioner in the building;

[0053] The cost model of building load coordinated regulation includes the total cost of system operation and user satisfaction loss. The sum of the cost of distributed generators, the cost of energy storage life loss, the cost of user satisfaction loss in the building caused by insufficient use of electric vehicle energy storage, and the cost of user satisfaction loss in the building caused by inappropriate indoor temperature is the total cost of system operation and user satisfaction loss:

[0054]

[0055] Among them, CALL (t) is the total cost.

[0056] Furthermore, the step S7 specifically includes:

[0057] The actual building load coordination problem is transformed into a deep reinforcement learning problem, and the components are defined, including the state space S, the action space A, the reward function r and the agent's strategy π. In the deep reinforcement learning algorithm, the agent defined by the algorithm interacts with the environment, learns the optimal strategy π and achieves specific goals through decision optimization under the deep reinforcement learning framework. The main function of the agent is to select the appropriate action a(t) according to the current state S(t), and continuously adjust its decision-making process through environmental feedback (reward or punishment) to maximize the expected cumulative reward R(t).

[0058] The state space S(t) contains all the information of the building system at time t, including the photovoltaic output P PV (t), rigid load P L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle arrival time Electric car departure time Indoor temperature Air conditioning on time Air conditioning off time It can be expressed as:

[0059]

[0060] However, some status information will leak the user's personal privacy information. For example, the arrival and departure time of electric vehicles can reflect the user's travel habits, the air conditioner on and off time can reflect whether there is someone in the building to control the switch, and the photovoltaic output and load will reveal the size of the building and the approximate number of users. For privacy reasons, some information is protected. To this end, the electric vehicle status and air conditioning status in 0 means that the air conditioner cannot participate in the regulation (the electric car is not at the charging station or the owner is unwilling to participate in the regulation, and the air conditioner does not participate in the indoor temperature regulation). 1 means that the air conditioner can participate in the regulation (the electric car participates in the regulation, and the air conditioner participates in the indoor temperature regulation). pure (t) is used to represent the difference between photovoltaic output and load, P pure (t) = P PV (t)-P L (t) is used to protect the photovoltaic output data and load data. In addition, the indoor temperature is hidden to protect the user's temperature preference.

[0061] Therefore, the state space is modified to:

[0062]

[0063] At this time, the state S′(t) contains all non-privacy protection information of the building system at time t, including the difference P between the photovoltaic output and the load L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status By defining an appropriate state space, the agent can better perceive changes in the building environment and respond accordingly;

[0064] The controllable devices include distributed generators and air conditioners in buildings. Therefore, the action space is expressed as:

[0065]

[0066] Among them, u DG (t) is the control signal of the distributed generator, is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signal of the load aggregator for the distributed generators and each air conditioner at time t;

[0067] The goal of the proposed building load coordinated control scheme is to minimize the operating cost of the entire system while ensuring user experience. In this sense, the reward at time t, that is, the higher the cost of system operation and user satisfaction loss, the smaller the reward. Therefore, the reward at time t is:

[0068]

[0069] At each time t, the agent obtains the building system state S′(t) and knows the difference P between the PV output and the load from the state S′(t). L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information, and select action a(t) from the action space A according to the strategy π, and obtain the control signals of the distributed generator and the air conditioner in the building from a(t), where π:Ω→A is a mapping from state S′(t) to action a(t), which describes the probability of the agent selecting action a(t) under a given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and the process continues until the terminal moment T. The discounted cumulative reward starting from moment t is defined as:

[0070]

[0071] Where γ is the discount factor, γ∈(0,1).

[0072] The goal of the deep reinforcement learning algorithm is to learn a strategy π that maximizes the initial reward J[π] = E[R(0)] maxJ[π]. Because the larger the reward r(t) at each step, the lower the system cost. Therefore, when the initial reward is the largest, the total cost of the entire building system operation process is the smallest.

[0073] Furthermore, the method for obtaining the final strategy π in step S8 is:

[0074] Use deep reinforcement learning methods to train the strategies π of all agents to maximize the objective function value J[π]. Strategy π describes the probability of the agent choosing action a(t) in a given state S′(t). By training strategy π, the agent chooses the action that maximizes the objective function in each state according to strategy π, thereby minimizing the operating cost of the entire system while ensuring user experience;

[0075] In order to better perform the strategy, define the value function V under the strategy π and the state value S′(t) π (S′(t))=E[R(t)|S′(t)]; where E[·] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under strategy π; the value function can be used to measure the quality of S′(t). If V π (S′(t)) is high, which means that starting from this state, the agent can obtain higher rewards by acting according to strategy π. In building load control, the value function V π (S′(t)) can be used to evaluate whether the current building state is beneficial to the goal. High-value states guide the agent to prioritize maintaining or transferring to more favorable building states (i.e., ensuring user experience while minimizing the operating cost of the entire system). Based on the value function, for each moment t, the advantage function is calculated The parameters θ and θ used to update the policy function and value function v ; θ and θ v is the policy function π(·|·; θ) and the value function V(·; θ v ), θ is the policy parameter. By optimizing θ, the policy π can learn how to choose the optimal action in different buildings and different states; θ v is the value function parameter, which is used to evaluate the overall performance under the current state (whether the building system operating cost is reduced and whether the user experience is guaranteed); parameters θ and θ v Use the following methods to update:

[0076]

[0077] in, represents the gradient of the policy function with respect to the parameter θ, which is used to adjust the probability of selecting an action; α and β are the parameters θ and θ respectively. v The learning rate determines the step size of each update;

[0078] By repeatedly updating the parameters θ v , improve the accuracy of state and action value estimation; based on the value function feedback, optimize the policy parameter θ to increase the probability of selecting high-advantage actions; repeat the above process until the output distribution of the policy function π(·|·; θ) remains almost unchanged in multiple consecutive iterations, indicating that the policy has stabilized and the value function V(·; θ v ) approaches zero, indicating that the value assessment is accurate enough; in the iterative process, the strategy π(·|·; θ) gradually tends to be able to select the action with the highest benefit in all states; as the iteration is completed, the load aggregator obtains the final strategy π. The load aggregator can use the obtained strategy π to calculate the building system state S′(t) (the difference between the photovoltaic output and the load P) at any time. L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status ), select the corresponding optimal action a(t) to control the distributed generators and air conditioners in the building, and realize the coordinated control of building loads considering the user needs of the V2B mode.

[0079] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0080] 1. This invention optimizes the charging and discharging strategy of electric vehicle batteries by improving the reward function in the deep reinforcement learning algorithm, thereby reasonably controlling the frequency and depth of charging and discharging, and reducing the loss of battery performance caused by excessive use. This method effectively extends the service life of energy storage equipment, reduces the economic cost caused by battery replacement, and thus improves user acceptance of the V2B model.

[0081] 2. The present invention designs a load control scheme that can protect user privacy. By improving the state matrix of the deep reinforcement learning algorithm, it effectively avoids the algorithm's excessive reliance on user sensitive data (electric vehicle arrival and departure time, photovoltaic output, building load, and air conditioning switching time in the building), thereby reducing the risk of privacy leakage and enhancing users' trust in the V2B technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 This is a schematic diagram of the implementation scenario of building load optimization and control under the V2B mode;

[0083] Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0084] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0085] like Figure 1 As shown, the present invention considers a building load optimization and control scheme based on deep reinforcement learning. In this scenario, the load aggregator located in the center interacts with all distributed buildings and resources through a deep reinforcement learning agent. The agent receives real-time data of the building (such as load demand, power generation of distributed energy, charging and discharging status of electric vehicles) and external signals (such as weather data) to form system state input. Each building unit has local energy scheduling capabilities (such as solar energy self-generation and self-use, air conditioning load adjustment, charging and discharging management of electric vehicles, etc.). Through deep reinforcement learning, the central control system can coordinate and optimize the load distribution of multiple buildings, dynamically allocate resources, and achieve energy balance and maximize utilization efficiency. In the figure, electric vehicles are connected to the building system through the V2B mode as mobile energy storage units. Deep reinforcement learning can predict the use behavior and battery status of electric vehicles, make charging and discharging decisions at the right time, thereby alleviating building load fluctuations and supporting grid operation. The deep reinforcement learning algorithm continuously updates the policy function through continuous environmental interaction (such as real-time building load data), so that the control model can adapt to the dynamic changes of building loads and grid conditions.

[0086] Based on the above implementation scenario, the present invention provides a building load collaborative control method considering the user needs of the V2B mode, such as Figure 2 As shown, it includes the following steps:

[0087] S1: In the coordinated control of building loads, random equations are used to simulate the power changes of building photovoltaics and rigid loads, and the power models of building photovoltaics and rigid loads are obtained after a series of feedback approximations;

[0088] Creating low-carbon buildings and using clean energy, mainly photovoltaic power generation, is an important measure for its green transformation. Building photovoltaics usually exist in a distributed form, which can generate electricity and consume it locally, reducing power transmission losses and improving energy efficiency. In buildings, there are some rigid loads that are not suitable for regulation, such as basic lighting loads, household electricity (kitchen refrigerators, microwave ovens, etc.) and office equipment (computers, printers, etc.). In the coordinated regulation of building loads, random equations are used to simulate the power changes of building photovoltaics and rigid loads. After a series of feedback approximations, the power model of building photovoltaics and rigid loads can be obtained as follows:

[0089]

[0090] Among them, t is the time when the system is in, P PV (t), P L (t) are the output power of building photovoltaic and rigid load at time t; P PVT (t) is the theoretical solar irradiance received by the photovoltaic at time t without considering weather conditions, is the overall trend of photovoltaic output power; ω PV (t) and ω L (t) is the standard Brownian motion; μ PV , μ L are the drift coefficients of building photovoltaics and rigid loads, σ PV , σ L are the diffusion coefficients of building PV and rigid loads, respectively;

[0091] μ PV , μ L , σ PV , σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows:

[0092]

[0093] Among them, ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, is the transition probability function;

[0094] By setting the partial differential derivative to 0, we can obtain the unknown system parameter μPV , σ PV , μ L , σ L Specific values:

[0095]

[0096] S2: Constructing mathematical model of distributed generators:

[0097]

[0098] Among them, P DG (t) is the generator output power at time t; T DG is the generator time constant, the specific value depends on the engine model; is the maximum output power of the generator; u DG (t) is the generator control signal.

[0099] S3: Constructing a mathematical model for electric vehicle battery energy storage:

[0100] V2B technology enables a two-way flow of energy between electric vehicles and buildings. Electric vehicles are charged when the building's electricity consumption is low, and power is supplied to the building during peak hours, which can smooth the load curve, consistent with the peak-shaving and valley-filling function of energy storage equipment.

[0101] The mathematical model of electric vehicle battery energy storage is:

[0102]

[0103] in, is the battery charge and discharge status of the i-th electric car at time t; P i BES (t) is the battery output power of the i-th electric vehicle at time t; Q s,i is the capacity factor of the i-th electric vehicle; η in,i , η out,i are the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle respectively.

[0104] S4: Constructing mathematical model of air conditioning load:

[0105] According to the purpose and characteristics, building loads include lighting loads, air conditioning and refrigeration loads, living and office equipment loads, etc. Among them, air conditioning is usually the largest single adjustable flexible load in a building, especially in commercial buildings and residences, and its power consumption accounts for 30%-50% of the total energy consumption. In addition, the energy consumption of air conditioning is affected by outdoor temperature, indoor comfort requirements, and user usage habits, showing strong volatility and flexibility. Therefore, the air conditioning load is modeled separately, and the mathematical model of the air conditioning load is:

[0106]

[0107] The change of indoor temperature is related to the outdoor temperature and the output power of the air conditioner. The mathematical model is as follows:

[0108]

[0109] in, is the indoor temperature of the jth room at time t; and is a coefficient determined by the corresponding room characteristics and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

[0110] S5: According to the model constructed in steps S1 to S4, establish a power balance constraint equation:

[0111]

[0112] Among them, N EV and N AC are the number of electric vehicles and air conditioners in the building scene, respectively.

[0113] S6: Based on the power balance constraint equation, a cost model for coordinated control of building loads is established. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of user satisfaction loss in the building due to insufficient demand for electric vehicle energy storage, and the cost of user satisfaction loss in the building due to inappropriate indoor temperature.

[0114] The mathematical expression of distributed generator cost is:

[0115]

[0116] Among them, C DG (t) is the generator cost at time t; and are the generator cost coefficients respectively;

[0117] By introducing the PLET model, the mathematical expression of the energy storage loss cost of electric vehicles is considered. This model expresses the energy storage life loss as LOH, The total life of the energy storage device is expressed as Where n is the total number of charges and discharges within the time range considered by the model; d i (t) is the depth of discharge in the i-th charge-discharge cycle, k p is the Peukert life constant, obtained by parameter identification, usually in the range of 1.1 to 1.3; for a specific d, can be regarded as a constant. Therefore, by limiting ΔCPLET , which can achieve the effect of extending the energy storage life.

[0118] Therefore, the energy storage life loss cost of the i-th electric vehicle at time t is expressed as c PLET,i (t):

[0119]

[0120] Considering the loss of user satisfaction in the building due to insufficient energy storage demand of electric vehicles, the cost is:

[0121]

[0122] in, is the satisfaction loss cost of the i-th electric car in the building at time t; and are the charging cost coefficients of the i-th electric vehicle in the building; The difference between the current power of the electric vehicle and the power expected by the user;

[0123] Consider the cost of lost building occupant satisfaction due to inappropriate indoor temperature:

[0124]

[0125] in, is the satisfaction loss cost of the jth air conditioner in the building at time t; T min,j is the minimum temperature limit set by the user in the room where the jth air conditioner is located in the building, T max,j is the maximum temperature limit set by the user in the room where the jth air conditioner is located in the building; and is the comfort loss weight coefficient of the jth air conditioner in the building;

[0126] The cost model of building load coordinated regulation includes the total cost of system operation and user satisfaction loss. The sum of the cost of distributed generators, the cost of energy storage life loss, the cost of user satisfaction loss in the building caused by insufficient use of electric vehicle energy storage, and the cost of user satisfaction loss in the building caused by inappropriate indoor temperature is the total cost of system operation and user satisfaction loss:

[0127]

[0128] Among them, C ALL (t) is the total cost.

[0129] S7: Based on the cost model of building load collaborative control, the actual building load collaborative control problem is transformed into a deep reinforcement learning problem;

[0130] The actual building load coordination problem is transformed into a deep reinforcement learning problem, and the components are defined, including the state space S, the action space A, the reward function r and the agent's strategy π. In the deep reinforcement learning algorithm, the agent defined by the algorithm interacts with the environment, learns the optimal strategy π and achieves specific goals through decision optimization under the deep reinforcement learning framework. The main function of the agent is to select the appropriate action a(t) according to the current state S(t), and continuously adjust its decision-making process through environmental feedback (reward or punishment) to maximize the expected cumulative reward R(t).

[0131] The state space S(t) contains all the information of the building system at time t, including the photovoltaic output P PV (t), rigid load P L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle arrival time Electric car departure time Indoor temperature Air conditioning on time Air conditioning off time It can be expressed as:

[0132]

[0133] However, some status information will leak the user's personal privacy information. For example, the arrival and departure time of electric vehicles can reflect the user's travel habits, the air conditioner on and off time can reflect whether there is someone in the building to control the switch, and the photovoltaic output and load will reveal the size of the building and the approximate number of users. For privacy reasons, some information is protected. To this end, the electric vehicle status and air conditioning status in 0 means that the air conditioner cannot participate in the regulation (the electric car is not at the charging station or the owner is unwilling to participate in the regulation, and the air conditioner does not participate in the indoor temperature regulation). 1 means that the air conditioner can participate in the regulation (the electric car participates in the regulation, and the air conditioner participates in the indoor temperature regulation). pure (t) is used to represent the difference between photovoltaic output and load, P pure (t) = P PV (t)-P L (t) is used to protect the photovoltaic output data and load data. In addition, the indoor temperature is hidden to protect the user's temperature preference.

[0134] Therefore, the state space is modified to:

[0135]

[0136] At this time, the state S′(t) contains all non-privacy protection information of the building system at time t, including the difference P between the photovoltaic output and the load L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status By defining an appropriate state space, the agent can better perceive changes in the building environment and respond accordingly;

[0137] The controllable devices include distributed generators and air conditioners in buildings. Therefore, the action space is expressed as:

[0138]

[0139] Among them, u DG (t) is the control signal of the distributed generator, is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signal of the load aggregator for the distributed generators and each air conditioner at time t;

[0140] The goal of the proposed building load coordinated control scheme is to minimize the operating cost of the entire system while ensuring user experience. In this sense, the reward at time t, that is, the higher the cost of system operation and user satisfaction loss, the smaller the reward. Therefore, the reward at time t is:

[0141]

[0142] At each time t, the agent obtains the building system state S′(t) and knows the difference P between the PV output and the load from the state S′(t). L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information, and select action a(t) from the action space A according to the strategy π, and obtain the control signals of the distributed generator and the air conditioner in the building from a(t), where π:Ω→A is a mapping from state S′(t) to action a(t), which describes the probability of the agent selecting action a(t) under a given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and the process continues until the terminal moment T. The discounted cumulative reward starting from moment t is defined as:

[0143]

[0144] Where γ is the discount factor, γ∈(0,1).

[0145] The goal of the deep reinforcement learning algorithm is to learn a strategy π that maximizes the initial reward J[π] = E[R(0)] maxJ[π]. Because the larger the reward r(t) at each step, the lower the system cost. Therefore, when the initial reward is the largest, the total cost of the entire building system operation process is the smallest.

[0146] S8: Iteratively solve the deep reinforcement learning problem to obtain the final strategy π. According to the final strategy π, the corresponding optimal action a(t) is selected to control the distributed generators and air conditioners in the building, so as to realize the coordinated control of building loads considering the needs of V2B mode users. Specifically, it includes:

[0147] The final strategy π is obtained as follows:

[0148] Use deep reinforcement learning methods to train the strategies π of all agents to maximize the objective function value Strategy π describes the probability of the agent choosing action a(t) in a given state S′(t). By training strategy π, the agent chooses the action that maximizes the objective function in each state according to strategy π, thereby minimizing the operating cost of the entire system while ensuring user experience;

[0149] In order to better perform the strategy, define the value function V under the strategy π and the state value S′(t) π (S′(t))=E[R(t)|S′(t)]; where E[·] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under strategy π; the value function can be used to measure the quality of S′(t). If V π (S′(t)) is high, which means that starting from this state, the agent can obtain higher rewards by acting according to strategy π. In building load control, the value function V π (S′(t)) can be used to evaluate whether the current building state is beneficial to the goal. High-value states guide the agent to prioritize maintaining or transferring to more favorable building states (i.e., ensuring user experience while minimizing the operating cost of the entire system). Based on the value function, for each moment t, the advantage function is calculated The parameters θ and θ used to update the policy function and value function v ; θ and θ v is the policy function π(·|·;θ) and the value function V(·;θ v), θ is the policy parameter. By optimizing θ, the policy π can learn how to choose the optimal action in different buildings and different states; θ v is the value function parameter, which is used to evaluate the overall performance under the current state (whether the building system operating cost is reduced and whether the user experience is guaranteed); parameters θ and θ v Use the following methods to update:

[0150]

[0151] in, represents the gradient of the policy function with respect to the parameter θ, which is used to adjust the probability of selecting an action; α and β are the parameters θ and θ respectively. v The learning rate determines the step size of each update;

[0152] By repeatedly updating the parameters θ v , improve the accuracy of state and action value estimation; based on the value function feedback, optimize the policy parameter θ to increase the probability of selecting high-advantage actions; repeat the above process until the output distribution of the policy function π(·|·; θ) remains almost unchanged in multiple consecutive iterations, indicating that the policy has stabilized and the value function V(·; θ v ) approaches zero, indicating that the value assessment is accurate enough; in the iterative process, the strategy π(·|·;θ) gradually tends to be able to select the action with the highest benefit in all states; as the iteration is completed, the load aggregator obtains the final strategy π. The load aggregator can use the obtained strategy π to calculate the building system state S′(t) (the difference between the photovoltaic output and the load P) at any time. L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status ), select the corresponding optimal action a(t) to control the distributed generators and air conditioners in the building, and realize the coordinated control of building loads considering the user needs of the V2B mode.

[0153] Based on the above scheme, in order to verify the actual effect of the method of the present invention, this embodiment experimentally compares the control method of the present invention with the existing control method, as follows:

[0154] After the experiment, the comparative data of battery life loss degree as shown in Table 1 was obtained.

[0155] Table 1

[0156]

[0157] It can be seen from Table 1 that the degree of battery life loss under the control method of the present invention is better than that of the other three methods, and the method of the present invention can prolong the service life of the battery.

Claims

1. A building load collaborative control method considering V2B mode user needs, characterized in that: The steps include: S1: In the coordinated control of building loads, random equations are used to simulate the power changes of building photovoltaics and rigid loads, and the power models of building photovoltaics and rigid loads are obtained after a series of feedback approximations; S2: Construct mathematical model of distributed generators; S3: Construct a mathematical model for electric vehicle battery energy storage; S4: Construct mathematical model of air conditioning load; S5: Establish a power balance constraint equation according to the model constructed in steps S1 to S4; S6: Based on the power balance constraint equation, a cost model for coordinated control of building loads is established. The cost model includes the cost of distributed generators, the depreciation cost of electric vehicle energy storage, the cost of user satisfaction loss in the building due to insufficient demand for electric vehicle energy storage, and the cost of user satisfaction loss in the building due to inappropriate indoor temperature. S7: Based on the cost model of building load collaborative control, the actual building load collaborative control problem is transformed into a deep reinforcement learning problem; S8: Iteratively solve the deep reinforcement learning problem to obtain the final strategy π. According to the final strategy π, the corresponding optimal action a(t) is selected to control the distributed generators and air conditioners in the building, so as to realize the coordinated control of building loads considering the user needs of the V2B mode.

2. According to claim 1, a building load collaborative control method considering V2B mode user needs is characterized in that: The power model of building photovoltaic and rigid load in step S1 is: P L (t)=μ L P L (t)dt+σ L dω L (t) (2) Among them, t is the time when the system is in, P PV (t), P L (t) are the output power of building photovoltaic and rigid load at time t; P PVT (t) is the theoretical solar irradiance received by the photovoltaic at time t without considering weather conditions, is the overall trend of photovoltaic output power; ω PV (t) and ω L (t) is the standard Brownian motion; μ PV , μ L are the drift coefficients of building photovoltaics and rigid loads, σ PV , σ L are the diffusion coefficients of building photovoltaics and rigid loads, respectively.

3. A building load collaborative control method considering V2B mode user needs according to claim 2, characterized in that: In step S1, μ PV , μ L , σ PV , σ L The four parameters are determined by the maximum likelihood estimation method, and the specific formula is as follows: Among them, ln(·) is the logarithm with the natural constant e as the base, L(·) is the likelihood function, is the transition probability function; By setting the partial differential derivative to 0, we can obtain the unknown system parameter μ PV , σ PV , μ L , σ L Specific values:

4. The method for collaborative control of building loads considering V2B mode user needs according to claim 1 is characterized in that: The mathematical model of the distributed generator in step S2 is: Among them, P DG (t) is the generator output power at time t; T DG is the generator time constant; is the maximum output power of the generator; u DG (t) is the generator control signal.

5. The method for collaboratively controlling building loads considering the needs of V2B mode users according to claim 1 is characterized in that: The electric vehicle battery energy storage mathematical model in step S3 is: in, is the battery charge and discharge status of the i-th electric car at time t; P i BES (t) is the battery output power of the i-th electric vehicle at time t; Q s,i is the capacity factor of the i-th electric vehicle; η in,i , η out,i are the charging efficiency coefficient and discharging efficiency coefficient of the i-th electric vehicle respectively.

6. The method for collaborative control of building loads considering V2B mode user needs according to claim 1 is characterized in that: The mathematical model of air conditioning load in step S4 is: The change of indoor temperature is related to the outdoor temperature and the output power of the air conditioner. The mathematical model is as follows: in, is the indoor temperature of the jth room at time t; and is a coefficient determined by the corresponding room characteristics and the thermal characteristics of the air conditioner; T OUT (t) is the outdoor temperature at time t.

7. The method for collaboratively controlling building loads considering V2B mode user needs according to claim 1 is characterized in that: The power balance constraint equation in step S5 is: Among them, N EV and N AC are the number of electric vehicles and air conditioners in the building scene, respectively.

8. The method for collaborative control of building loads considering V2B mode user needs according to claim 1 is characterized in that: The mathematical expression of the distributed generator cost in step S6 is: Among them, C DG (t) is the generator cost at time t; and are the generator cost coefficients respectively; By introducing the PLET model, the energy storage life loss cost of the i-th electric vehicle at time t is expressed as c PLET,i (t): Considering the loss of user satisfaction in the building due to insufficient energy storage demand of electric vehicles, the cost is: in, is the satisfaction loss cost of the i-th electric car in the building at time t; and are the charging cost coefficients of the i-th electric vehicle in the building; The difference between the current power of the electric vehicle and the power expected by the user; Consider the cost of lost building occupant satisfaction due to inappropriate indoor temperature: in, is the satisfaction loss cost of the jth air conditioner in the building at time t; T min,j is the minimum temperature limit set by the user in the room where the jth air conditioner is located in the building, T max,j is the maximum temperature limit set by the user in the room where the jth air conditioner is located in the building; and is the comfort loss weight coefficient of the jth air conditioner in the building; The cost model of building load coordination includes the total cost of system operation and user satisfaction loss: Among them, C ALL (t) is the total cost.

9. A building load collaborative control method considering V2B mode user needs according to claim 8, characterized in that: The step S7 specifically includes: Define the components, including state space S, action space A, reward function r and agent strategy π; in the deep reinforcement learning algorithm, the agent defined by the algorithm interacts with the environment, learns the optimal strategy π and achieves specific goals through decision optimization under the deep reinforcement learning framework; the agent selects the appropriate action a(t) according to the current state S(t), and continuously adjusts its decision-making process through environmental feedback to maximize the expected cumulative reward R(t); among them, The state space S(t) is expressed as: Modify the state space to: At this time, the state S′(t) contains all non-privacy protection information of the building system at time t, including the difference P between the photovoltaic output and the load L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status The action space is represented as: Among them, u DG (t) is the control signal of the distributed generator, is the air conditioning control signal; the action space a(t) of the deep reinforcement learning algorithm represents the control signal of the load aggregator for the distributed generators and each air conditioner at time t; The reward at time t is: At each time t, the agent obtains the building system state S′(t) and knows the difference P between the PV output and the load from the state S′(t). L (t), distributed generator output power P DG (t), external temperature T OUT (t), electric vehicle energy storage charging and discharging status Electric vehicle status and air conditioning status These state information, and select action a(t) from the action space A according to the strategy π, and obtain the control signals of the distributed generator and the air conditioner in the building from a(t), where π:Ω→A is a mapping from state S′(t) to action a(t), which describes the probability of the agent selecting action a(t) under a given state S′(t); in return, the agent receives a reward r(t) and the next state S′(t+1), and the process continues until the terminal moment T. The discounted cumulative reward starting from moment t is defined as: Where γ is the discount factor, γ∈(0,1).

10. The method for collaborative control of building loads considering V2B mode user needs according to claim 8, characterized in that: The method for obtaining the final strategy π in step S8 is: Use deep reinforcement learning methods to train the strategies π of all agents to maximize the objective function value J[π]. Strategy π describes the probability of the agent choosing action a(t) in a given state S′(t). By training strategy σ, the agent chooses the action that maximizes the objective function in each state according to strategy π, thereby minimizing the operating cost of the entire system while ensuring user experience. In order to better perform the strategy, define the value function V under the strategy π and the state value S′(t) π (S′(t))=E[R(t)|S′(t)]; where E[·] represents the mathematical expectation; the value function V π (S′(t))=E[R(t)|S′(t)] represents the expected reward accumulated from the current state S′(t) to all future time steps under strategy π; based on the value function, for each time t, the advantage function is calculated The parameters θ and θ used to update the policy function and value function v ; parameters θ and θ v Use the following methods to update: in, represents the gradient of the policy function with respect to the parameter θ, which is used to adjust the probability of selecting an action; α and β are the parameters θ and θ respectively. v The learning rate determines the step size of each update; By repeatedly updating the parameters θ v , improve the accuracy of state and action value estimation; based on the value function feedback, optimize the strategy parameter θ to increase the probability of selecting high-advantage actions; as the iteration is completed, the load aggregator obtains the final strategy π.

Citation Information

Patent Citations

  • Joint computing unloading and resource allocation method based on multi-agent DDQN

    CN114584951A

  • Power distribution network layered and partitioned load shedding coordination control method based on event trigger mechanism

    CN114844050A

  • Networked electric vehicle cooperative charging and discharging regulation and control method based on game between users

    CN115626072A

  • Intelligent building group electricity-carbon joint distributed transaction strategy acquisition method and device

    CN115809568A

  • Intelligent building electricity demand side optimization scheduling strategy based on particle swarm optimization

    CN116579571A