A method and device for energy storage system management based on deep reinforcement learning
By employing a two-layer optimization strategy based on deep reinforcement learning, combining CEEMDAN and extreme learning machine to extract load features, and using an improved beetle algorithm to solve the optimal control strategy, the problems of personalized electricity consumption strategies and load monitoring and identification errors in residential intelligent energy storage management systems have been solved, resulting in higher economic benefits and user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CYG SUNRI CO LTD
- Filing Date
- 2022-03-15
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies in residential smart energy storage management systems lack personalized electricity consumption strategies for different users, have large load monitoring and identification errors, and heuristic optimization algorithms are prone to getting trapped in local optima, making it difficult to balance economic benefits and user satisfaction.
A two-layer optimization strategy based on deep reinforcement learning is adopted. By constructing an energy storage system model and safety constraints, load features are extracted by combining CEEMDAN and extreme learning machine. An improved beetle algorithm is used to solve the optimal control strategy. An action-penalty integrated function is constructed to optimize the economic benefits and user satisfaction of the energy storage system.
It has achieved higher economic benefits and user satisfaction, improved the accuracy of load identification and the optimization effect of control strategies, and enhanced the overall performance of energy storage systems.
Smart Images

Figure CN114744651B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of deep learning, and particularly relates to a method and device for managing an energy storage system based on deep reinforcement learning. BACKGROUND
[0002] The large-scale access of distributed energy sources causes a sharp increase in uncertain factors of microgrids, and household energy storage systems also face great challenges. For example, energy supply response is not timely, energy utilization rate is low, and the like. Based on an intelligent energy storage system, a reliable, safe, and efficient energy management system can not only realize real-time management and control of household electrical equipment, but also greatly improve the reliability of the microgrid system. Thus, the capacity credibility of large-scale grid connection of distributed energy sources is improved, the power quality of the grid is optimized, and the dispatching difficulty is reduced.
[0003] At present, there are not many studies on household intelligent energy storage management systems at home and abroad. At present, most of them are concentrated in two aspects of energy management systems and load monitoring. The current control strategy of the energy management system is mainly divided into grid-connected type, off-grid type, and off-grid and grid-connected integrated type, and the focus is on peak clipping, fluctuation suppression, and the like. However, these strategies are too general, and different power strategies are not formulated for different users. The load monitoring stays in the direction of identification, and most of them adopt the invasive and data collection methods to identify the electrical load. However, the error of the identification result is too large, the credibility is not high, and the result has no practical effect. Finally, for the intelligent optimization algorithm, the heuristic optimization algorithm is mostly used for model solving. However, the heuristic optimization algorithm is prone to local optimum, and the convergence precision and convergence speed are limited by itself, it is difficult to judge the overall load, and it is difficult to consider both economic benefits and user satisfaction. SUMMARY
[0004] In a first aspect, the application provides a method for managing an energy storage system based on deep reinforcement learning, which comprises the following steps:
[0005] establishing an energy storage system model;
[0006] setting a safety constraint condition according to the energy storage system model;
[0007] constructing a double-layer state space of the energy storage system meeting the safety constraint condition, wherein the double-layer state space comprises a first-layer state space for representing optimal scheduling of maximum economic benefit and highest satisfaction of the energy storage system; and a second-layer state space for representing charge and discharge control of the energy storage system;
[0008] constructing a double-layer action space corresponding to the double-layer state space and obtaining a feedback value of the double-layer state space generated by the control instruction when the energy storage system model receives the control instruction, wherein the double-layer action space comprises a first-layer action space for representing a variation of system load; and a second-layer action space for representing a variation of charging and discharging;
[0009] constructing an action-punishment integrated function in combination with the double-layer action space, the action-punishment integrated function being used for giving the energy storage system model positive or negative punishment according to the feedback value;
[0010] inference of a control strategy through the action-punishment integrated function to obtain a control strategy with the highest economic benefit and user satisfaction of the energy storage system.
[0011] Further, the energy storage system model is obtained by the following formula:
[0012]
[0013] wherein, P ESS is the rated power of the energy storage system; Q ESS is the rated capacity of the energy storage system; represents the charging efficiency of the energy storage system, represents the discharging efficiency of the energy storage system; represents the state of charge variable of the energy storage system at t time; represents the state of discharge variable of the energy storage system at t time.
[0014] Further, the safety constraint condition comprises a battery power constraint, an energy storage system periodic constraint, an energy storage system charging and discharging constraint, an energy storage system total power constraint, and a peak load shifting constraint, wherein,
[0015] the battery power constraint is represented by the formula , and respectively represent the upper and lower limits of the state of charge of the energy storage system;
[0016] the energy storage system periodic constraint is represented by the formula , represents the remaining storage power of the energy storage system after the end of a period, represents the initial set power of the energy storage system at the beginning of the next operation period;
[0017] the energy storage system charging and discharging times constraint is represented by the formula , end represents the single period operation time, represents the charging and discharging activation times of the energy storage system, N ESSrepresents the maximum number of charge and discharge activations of the energy storage system in a single operating cycle;
[0018] The total power constraint of the energy storage system scheduling is represented by the formula represents, represents the total load power of the bottom layer of the double-layer control strategy at time t; max represents the maximum power allowed by the energy storage system;
[0019] The peak clipping constraint is represented by the formula out.t ≤F put.max represents, P out.t represents the real-time output power of the energy storage system, F out.max represents the peak load in a cycle period.
[0020] Further, the state equation of the first layer state space is: where θ t represents the real-time electricity price of the state space at time t, represents the distributed energy output at time t, T represents the system load mobilization time, E t-1 EV represents the state of charge of the electric vehicle at the previous time;
[0021] The state equation of the second layer state space is: where E t-1 ESS represents the total load power transmitted from the first layer state equation to the second layer state equation at time t.
[0022] Further, the action equation of the first layer action space is: where, represents the rigid load, represents the time-variable load, represents the power-variable load, represents the electric vehicle charging load;
[0023] The action equation of the second layer action space is:
[0024] Further, when the energy storage system model receives a control instruction, the feedback value generated by the double-layer state space due to the control instruction includes:
[0025] The CEEMDAN combined with the extreme learning machine is used to extract the decomposed load characteristics from the original load curve and to confirm the load type, which includes rigid load, time-variable load, power-variable load and electric vehicle charging load.
[0026] Outputting waveforms of various load types, solving the double-layer action space to obtain feedback values.
[0027] Further, the outputting waveforms of various load types, solving the double-layer action space to obtain feedback values, comprises:
[0028] Solving the double-layer action space by using a genetic algorithm.
[0029] Further, the constructing an action-penalty integrated function in combination with the double-layer action space comprises the following steps:
[0030] Mapping the optimization objective function of the load of the energy storage system model and the photovoltaic output into a first-layer penalty function;
[0031] Taking the difference between the charging and discharging benefits and the over-limit penalty cost of the energy storage system as a second-layer penalty function;
[0032] Presetting state quantities and scheduling quantities, and evaluating the first-layer penalty function and the second-layer penalty function to obtain an optimal action-penalty integrated function according to an expected sum of penalties achieved at preset time steps.
[0033] Further, the first-layer penalty function is: wherein represents a penalty for a charging behavior violating operation constraints, and the second-layer penalty function is: wherein represents charging and discharging benefits, is an over-limit penalty cost.
[0034] In a second aspect of the present application, a household energy storage system management device is provided, comprising:
[0035] An energy storage system model construction module is configured to construct an energy storage system model.
[0036] A safety constraint condition setting module is configured to set safety constraint conditions for the energy storage system model, and the safety constraint conditions are used to ensure safe and stable operation of the energy storage system model.
[0037] A double-layer state space construction module is configured to construct a double-layer state space of the energy storage system model in compliance with the safety constraint conditions, and the double-layer state space is used to represent an external environment in which the energy storage system model is located.
[0038] A double-layer dynamic space construction module is configured to construct a double-layer dynamic space corresponding to the double-layer state space, and the double-layer dynamic space is used to embody changes of the double-layer state space after the energy storage system model receives a control instruction.
[0039] The learning module is configured to set up an action-punishment integrated function, and learn the control instruction strategy through forward and backward punishment.
[0040] The control instruction strategy reasoning module is configured to reason a control instruction strategy with the highest economic benefit and user satisfaction.
[0041] In a third aspect, the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the computer program.
[0042] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement the method described above.
[0043] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0044] The present application proposes a hierarchical optimization strategy based on deep learning for intelligent energy storage systems, which has better economic benefits, higher user satisfaction, and strong practicability.
[0045] The present application introduces the CEEMDAD algorithm, extracts and models different power consumption characteristics of user-side loads according to the decomposition strategy of the algorithm, and realizes load identification in combination with extreme learning machines. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0047] Figure 1 is a flowchart of a deep reinforcement learning-based energy storage system management method provided by an embodiment of the present application;
[0048] Figure 2 is a schematic diagram of a double-layer optimization strategy principle used in a deep reinforcement learning-based energy storage system management method provided by an embodiment of the present application;
[0049] Figure 3 is a schematic diagram of decomposition of user loads by a CEEMDAN load extraction method provided by an embodiment of the present application;
[0050] Figure 4is a flowchart of an improved ant algorithm provided by an embodiment of the present application. DETAILED DESCRIPTION
[0051] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.
[0052] The technical solutions adopted by the present application are further described below in combination with the drawings and embodiments.
[0053] Embodiment 1
[0054] Referring to Figure 1 A deep reinforcement learning-based energy storage system management method provided by the present application includes the following steps:
[0055] S1: Establish an energy storage system model. A household energy storage system generally uses a battery as an energy storage element, and the charging and discharging process of the energy storage system can be controlled through a pre-established optimization strategy. The charging and discharging power of the energy storage system at time t is and the real-time battery power is which can be expressed by formula (1):
[0056]
[0057] In the above formula: P ESS is the rated power of the energy storage system; Q ESS is the rated capacity of the energy storage system; represents the charging efficiency of the energy storage system, represents the discharging efficiency of the energy storage system; v t.ch ESS represents the state of charge variable of the energy storage system at time t; represents the state of discharge variable of the energy storage system at time t; In particular, when the state of charge variable or the state of discharge variable is 1, it means that the energy storage system is in a working state, and when the state of charge variable or the state of discharge variable is 0, the energy storage system is in a standby state. These two states cannot exist at the same time.
[0058] S2: Set safety constraints. In order to enable the energy storage system to operate better and more safely, operating constraints should be set for the energy storage system according to different situations;
[0059] For example, the battery power constraint of the energy storage system can be expressed by formula (2):
[0060]
[0061] In the above formula, E min ESS respectively represent the upper and lower limits of the state of charge of the energy storage system.
[0062] In order to enable the energy storage system to have a more efficient control mode and be able to operate for a long time, the periodic constraint of the energy storage system can be represented by formula (3):
[0063]
[0064] In the above formula, represents the remaining storage power of the energy storage system after a period, represents the initial setting power of the energy storage system at the beginning of the next operating period, ensuring that the power at the end of each period and at the beginning of the next period is the same.
[0065] In order to reduce the number of charge and discharge activations of the energy storage system, thereby improving the service life of the energy storage system, the charge and discharge frequency constraint of the energy storage system can be represented by formula (4):
[0066]
[0067] In the above formula, T end represents the single period operating time, represents the number of charge and discharge activations of the energy storage system. N ESS represents the maximum number of charge and discharge activations of the energy storage system per operating period.
[0068] In order to enable the load of the energy storage system to concentrate even in the low price valley stage without exceeding the low voltage limit, the total power of the energy storage system should be lower than the set threshold, so the total power constraint of the energy storage system can be represented by formula (5):
[0069] P t 1 +P t ESS ≤P max (5)
[0070] In the above formula, P t 1 represents the total load power of the bottom layer of the double-layer control strategy at time t; P max represents the maximum power allowed by the energy storage system.
[0071] In order to maintain the stability of the microgrid, the energy storage system should have a peak clipping and valley filling constraint, which can be represented by formula (6):
[0072] Pout.t ≤F out.max (6)
[0073] In the above formula, P out.t represents the real-time output power of the energy storage system. F out.max represents the peak load in a cycle period.
[0074] S3: Set the intelligent energy management system
[0075] As Figure 2 shown, the intelligent energy management system can be simply regarded as a mathematical optimization process, with the change of time t, the energy management system can change to make timely response, and when the system executes the current instruction, it can continue to receive the instruction brought by t+1 time, until the end of the period, based on this, a double-layer optimization model is established.
[0076] (1) Establish the state space
[0077] The state space here refers to the external environment of the energy storage system, the more detailed the state space expression is, the more regular the energy storage system changes to the external environment.
[0078] For example, the first layer state space involves the optimization scheduling of the target function of maximizing the economic benefit of energy storage and the highest satisfaction, and its state space expression is shown in formula (7).
[0079]
[0080] In the above formula, θ t represents the real-time electricity price of the state space at t time, represents the distributed energy output at t time, T represents the system load moving time, and E t-1 EV represents the state of charge of the electric vehicle charging at the last time.
[0081] The second layer state space, on the basis of the first layer, involves the charge and discharge control of the energy storage system, and its state space expression is shown in formula (8).
[0082]
[0083] In the above formula, E t-1 ESS represents the total load power of the first layer state space transmitted to the second layer state space at t time.
[0084] (2) Establish the action space
[0085] The action space here refers to a reaction to external changes after the intelligent energy storage system receives instructions. The more detailed the action space expression, the clearer the instructions.
[0086] First layer action space: including various loads of the energy storage system, the action space expression is shown in equation (9):
[0087]
[0088] Here P k.t Fixed ,V m.t Time , is the set of system loads, including four, respectively representing rigid load, time variable load, power variable load, and electric vehicle charging load.
[0089] Second layer action space: execute the plan composed of charge and discharge state variables with the goal of maximizing long-term benefits. The action space expression is shown in equation (10):
[0090]
[0091] S4: Construct a penalty function. In combination with step S3, after reflecting the external changes, the system model will give the energy storage system positive or negative punishment according to the feedback value. In deep learning, it will tend to positive punishment and avoid negative punishment.
[0092] First layer: map the optimization objective functions of the four system loads and the photovoltaic output to the penalty function. Then the first layer penalty function at time t is shown in equation 11:
[0093]
[0094] In the above equation, represents the penalty for charging behavior violating operating constraints, which is expanded as shown in equation 12:
[0095]
[0096] In the above equation, and β t represent the penalty coefficients of excessive charging and excessive discharging respectively. δ t represents the penalty coefficient of user preference.
[0097] Second layer penalty coefficient: mainly includes the charging and discharging benefits of the energy storage system and the penalty cost of exceeding the limit, as shown in equation (13).
[0098]
[0099] In the above equation, represents the charging and discharging benefit, a running over-limit penalty fee. is expanded as shown in equation (14).
[0100]
[0101] In the above equation, represents the fee for violating the energy storage system battery constraint, and the specific equation is shown in equation 15:
[0102]
[0103] In the above equation, μ t respectively represent the penalty coefficients for overcharging and overdischarging of the intelligent energy storage system, and σ t represents the penalty coefficient for the energy storage system not reaching the set value after the end of a control action, is the end time of the energy storage system control.
[0104] represents the fee for violating the energy storage system charging and discharging conversion times, and the specific equation is shown in equation (16):
[0105]
[0106] In the above equation, ρ t represents the penalty coefficient for the energy storage system charging and discharging times exceeding the maximum conversion times after the end of a control action.
[0107] represents the fee for violating the total power over-limit, and the specific equation is shown in equation (17):
[0108]
[0109] In the above equation, zzx represents the penalty coefficient for the total power over-limit at time t.
[0110] represents the fee for violating the peak clipping and valley filling constraint, and the specific equation is shown in equation (18):
[0111]
[0112] In the above equation, λ t represents the penalty coefficient for the peak clipping and valley filling over-limit at time t.
[0113] S5, establishing an action-penalty integrated function to obtain an optimal action penalty function
[0114] The state quantity is set in advance, and its scheduling quantity a t The penalty r tthe expected sum, Ut, is evaluated as shown in equation (19):
[0115] U t = r t + γr t+1 + γ 2 r t+1 +... + γr K-t = r K + U t + U t+1 (19)
[0116] The Bellman equation for this is shown in equation (20):
[0117] Q π (s t , a t ) = E π [ U t | S t = s t , A t = a t ] = E π [ r t + γQ π (s t+1 , a t+1 ) | S t = s t , A t = a t ] (20)
[0118] In the above equation, E π denotes the total expected value. S t and A t denote the state and action variables of the input required by the system at time t, respectively. s t and a t denote the state and action variables received by the system at time t, respectively. π denotes the coefficient of the energy storage system mapping to the energy dispatching strategy. γ ∈ [0, 1] denotes the discount rate of the next step penalty to the current penalty. It should be noted that when γ is 1, the next step penalty is as important as the current penalty, and the intelligent energy storage system has foresight, while when γ is 0, the next step penalty does not interfere with the current variable, and the intelligent energy storage system is shortsighted. The ultimate goal in the dispatching process is to seek the optimal strategy and obtain the optimal action penalty function, as shown in equation 21.
[0119]
[0120] Example 2
[0121] Based on the above embodiment 1, the first layer action space extracts load characteristics, which can be extracted by CEEMDAN, and the extraction method is as follows:
[0122] Step one: use CEEMDAN to decompose the original load curve P(t)+ε0ω t (t) is a Gaussian white noise with positive integer amplitude) for M times of experiments, and use EMD algorithm for decomposition, and obtain the first modal component I i,1 , the first component decomposed by CEEMDAN is the mean value of all I i,1 of M times of experiments, as shown in formula (22):
[0123]
[0124] Step two: in step one, calculate the first residual function r1(t), as shown in formula (23):
[0125]
[0126] Step three: for the sequence r1(t)+ε1ω t (t) (wherein, ε1 is a Gaussian white noise constant, E1 is an EMD decomposition operation value), M times of modal decomposition are performed to obtain the first intrinsic modal component, and at this time, the second modal component 22 can be obtained according to the current residual. As shown in formula (24):
[0127]
[0128] Step four: calculate the remaining stage k, repeat steps 1-3, and calculate the k-1th intrinsic modal component according to formula (25) and (26).
[0129]
[0130] Step five: perform step 4 until the residual signal is less than a threshold value, and the standard is that the number of extreme points is not more than 2. Finally, the number of all intrinsic modal components is K, and the residual signal is as shown in formula (27).
[0131]
[0132] Step six: compare the intrinsic modal component with the load characteristics, extract the decomposition characteristics combined with the extreme learning machine, and confirm the load type. As shown in formula (28), it is mainly divided into four kinds of loads: rigid load, time variable load, power variable load and electric vehicle charging load. Figure 3
[0133] Among them, the rigid load mainly includes monitoring, smart meter and other equipment, and the reliability requirement is higher. Such power supply needs to run at rated power, and the decomposed load characteristic is the sum of the rated power of all rigid loads within a certain time.
[0134] The time-variable load mainly includes washing machines, rice cookers, oil smoke exhausters and other electrical equipment with adjustable running time. Such load can be turned on or off at a set time, and cannot be interrupted once started. The decomposed load characteristic suddenly increases or decreases within a certain time and remains for a long time.
[0135] The power-variable load mainly includes air conditioners, water heaters and other electrical equipment whose power can be controlled in multiple stages. Such load can adjust the gear within the power range, and the power presents a curve link when the gear changes. The decomposed load characteristic presents a straight line change in power at a certain time, and the straight line change presents a curve link.
[0136] Electric vehicle charging load, because the charging time and charging amplitude of electric vehicles are uncontrollable. Therefore, all irregular loads after decomposition are calculated as charging loads. Such load is calculated and modeled by considering the arrival time of the user, the initial state of charge, the expected state of charge when leaving, and the expected time of leaving. The use of the battery is considered comprehensively, and the battery life model is modeled around the battery.
[0137] Step seven: output waveform, which is used to solve the action space model to obtain feedback value.
[0138] Embodiment 3
[0139] The application also provides an improved bumblebee algorithm for model solving in the application, and the specific implementation process is as shown in Figure 4 .
[0140] Referring to Figure 4 , first, population initialization is performed, and the number of search agents, i.e. the number of bumblebees, is set. The population code is C, and in the S-dimensional (the number of targets to be solved) space, the position of the i-th bumblebee represents the optimal feasible solution in the space at the current iteration number. And set the maximum iteration number T max , the upper limit and lower limit of the search space.
[0141] After that, the current fitness value is calculated, and the current optimal solution is selected. In solving the double-layer optimization problem of household intelligent energy storage systems, the best fitness value of each iteration represents the value of the optimal solution of the current iteration, that is, the optimal economic benefit value and user satisfaction value sought by the model. The best fitness value composed of the termites represents the current optimal feasible solution, that is, the household storage system capacity configuration, power size, and user scheduling amount. The optimal solution and the optimal feasible solution are randomly selected in the first iteration.
[0142] Then, the position of the termite is updated. However, in the iteration process of the heuristic optimization algorithm, the iteration is performed in a specified space to update the position. However, this limits the population diversity and easily falls into a local optimum. In this embodiment, a Tend chaotic search mechanism is introduced, as shown in formula (28). The previous plane search mechanism is mapped to the space, and the position updating mechanism is transformed on the basis of ensuring the original population.
[0143]
[0144] In the above formula (28), represents the i-th search agent in the j-th dimensional space. T t represents the current iteration number.
[0145] Then, the Levy flight strategy is added to improve the search space. In the termite algorithm, the step size mechanism is the key to affecting the search ability of the algorithm. The initial step size should be set to be relatively large, so as to easily cover the entire search area and avoid falling into a local optimum. With the increase of the iteration number, the step size value needs to be gradually reduced, so as to facilitate the improvement of the search precision. Here, an adaptive flight strategy updating mechanism is introduced to automatically update the step size value.
[0146] Then, the adaptive weight is added, and the boundary repair is performed.
[0147] After multiple iterations, if the current population fitness value is better than the historical fitness value, the current population optimal solution and the feasible solution are updated, and the current optimal solution is used to replace the historical optimal solution. The value of the feasible solution of the current optimal solution is retained. The difference between the current optimal solution and the historical optimal solution is calculated.
[0148] When the maximum iteration number is reached or the difference between the optimal solution and the historical optimal solution is continuously less than a certain threshold value, the iteration is terminated. It should be noted that in the solution of a complex model, it is generally selected to judge whether the maximum iteration number is reached, and in the test algorithm, it is generally selected to judge whether the difference between the optimal solution and the historical optimal solution is continuously less than a certain threshold value.
[0149] Finally, the value of the model optimal solution and the value of the solution represented thereby are output.
[0150] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A management method for energy storage systems based on deep reinforcement learning, characterized in that, The method includes: Establish an energy storage system model; Safety constraints are set according to the energy storage system model; A two-layer state space for the energy storage system that meets the aforementioned safety constraints is constructed, wherein the two-layer state space includes a first-layer state space for representing the optimal scheduling that maximizes the economic benefits and satisfaction of the energy storage system; and a second-layer state space for representing the charging and discharging control of the energy storage system. Construct a two-layer action space corresponding to the two-layer state space and obtain the feedback value generated by the two-layer state space due to the control command when the energy storage system model receives the control command. The two-layer action space includes a first-layer action space for representing the change in system load and a second-layer action space for representing the change in charging and discharging. An integrated action-penalty function is constructed by combining the two-layer action space. This integrated action-penalty function is used to impose positive or negative penalties on the energy storage system model based on the feedback value. By employing an action-penalty integrated function reasoning control strategy, we aim to obtain the control strategy that maximizes the economic benefits and user satisfaction of the energy storage system. The energy storage system model is obtained through the following formula: ; in, For the energy storage system in t The charging and discharging power at any given moment; For the energy storage system in t Real-time battery level; This represents the total load power transferred from the first layer state space to the second layer state space at time t; The interval from time t-1 to time t; P ESS This refers to the rated power of the energy storage system. Q ESS This refers to the rated capacity of the energy storage system. This indicates the charging efficiency of the energy storage system. Indicates the discharge efficiency of the energy storage system; This represents the charging state variable of the energy storage system at time t; This represents the discharge state variable of the energy storage system at time t.
2. The method as described in claim 1, characterized in that, The safety constraints include battery capacity constraints, energy storage system periodicity constraints, energy storage system charge / discharge constraints, total power constraints for energy storage system scheduling, and peak shaving / valley filling constraints. Battery power constraint is expressed as follows express, and These represent the upper and lower limits of the state of charge of the energy storage system, respectively. The real-time battery charge of the energy storage system at time t-1; Periodic constraints of energy storage systems are expressed in the form of formula express, This indicates the remaining stored energy in the energy storage system after one cycle. This indicates the initial set capacity of the energy storage system at the start of the next operating cycle; The energy storage system charge / discharge cycle constraint is expressed as follows: express, Indicates the runtime of a single cycle. Indicates the number of charge-discharge activations of the energy storage system. This indicates the maximum number of charge / discharge activations of the energy storage system within a single operating cycle; The total power constraint for energy storage system dispatch is expressed as follows: express, This represents the total load power of the bottom layer of the two-layer control strategy at time t; This indicates the maximum allowable power of the energy storage system; For the energy storage system in t The charging and discharging power at any given moment; Peak shaving and valley filling constraints are based on formula express, This represents the real-time output power of the energy storage system. It is represented as the peak load within one cycle.
3. The method as described in claim 2, characterized in that, The state equations of the first layer state space are: ,in, This is the first layer of state space; Let be the real-time electricity price in the state space at time t. P t PV This represents the distributed energy output at time t in the state space, where T represents the system load shunting time. E t-1 EV This indicates the state of charge of the electric vehicle at the previous moment; The state equations of the second-level state space are: ,in, This is the second layer of state space; This represents the total load power transferred from the first layer state space to the second layer state space at time t.
4. The method as described in claim 3, characterized in that, The motion equations for the first layer of motion space are: ,in, This represents the first layer of action space; Indicates rigid load. Indicates time-variable load, P n.t Power Indicates variable power load ,V l.t EV Indicates the charging load of electric vehicles; The motion equations for the second-level motion space are: ,in, This represents the second layer of action space.
5. The method as described in claim 4, characterized in that, The step of obtaining the feedback value generated by the control command in the two-layer state space when the energy storage system model receives the control command includes: The CEEMDAN combined with Extreme Learning Machine was used to extract decomposed load features from the original load curve and identify the load types, including rigid loads, time-variable loads, power-variable loads, and electric vehicle charging loads. Output waveforms for each load type and solve the double-layer action space to obtain feedback values.
6. The method as described in claim 5, characterized in that, The output waveforms for each load type are used to solve the two-layer action space to obtain feedback values, including: The two-layer action space is solved using the beetle algorithm.
7. The method as described in claim 6, characterized in that, The action-penalty integrated function is constructed by combining the two-layer action space. Includes the following steps: The objective function for optimizing the load and the photovoltaic output of the energy storage system model are mapped to the first-layer penalty function; The difference between the charging and discharging revenue of the energy storage system and the over-limit penalty cost is used as the second-level penalty function; The first-layer penalty function and the second-layer penalty function are evaluated by setting the state quantity and the scheduling quantity, and the expected sum of the penalties achieved at the preset time step to obtain the optimal action-penalty integrated function.
8. The method as described in claim 7, characterized in that, The first-level penalty function is: ,in C l.t EV This indicates a penalty for violating operational constraints during charging. This is represented as the real-time electricity price in the state space at time t; This indicates the photovoltaic output; This represents the objective function for optimizing the rigid load. This represents the optimization objective function for the time-variable load. This represents the optimization objective function for the variable power load. The objective function representing the optimization of electric vehicle charging load; This represents the first-level penalty function; The second-level penalty function is: ,in Z t ESS Indicates the benefits of charging and discharging. C t ESS Penalties will be charged for exceeding the limits. This represents the second-level penalty function.
9. A residential energy storage system management device, characterized in that, include: An energy storage system model building module is used to build an energy storage system model, which is obtained through the following formula: ; in, For the energy storage system in t The charging and discharging power at any given moment; For the energy storage system in t Real-time battery level; This represents the total load power transferred from the first-level state space to the second-level state space at time t; The interval from time t-1 to time t; P ESS This refers to the rated power of the energy storage system. Q ESS This refers to the rated capacity of the energy storage system. This indicates the charging efficiency of the energy storage system. Indicates the discharge efficiency of the energy storage system; This represents the charging state variable of the energy storage system at time t; This represents the discharge state variable of the energy storage system at time t; A safety constraint setting module is used to set safety constraints for the energy storage system model, and the safety constraints are used to ensure the safe and stable operation of the energy storage system model. A two-layer state space construction module is used to construct a two-layer state space for an energy storage system model that meets safety constraints. The two-layer state space is used to represent the external environment in which the energy storage system model is located. A dual-layer dynamic space construction module is used to construct a dual-layer dynamic space corresponding to the dual-layer state space. The dual-layer dynamic space is used to reflect the changes in the dual-layer state space after the energy storage system model receives control commands. The learning module is used to establish an integrated action-punishment function, which learns the control instruction strategy through positive and negative punishment. The control instruction strategy reasoning module is used to reason about the control instruction strategy that yields the highest economic benefits and user satisfaction.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
KR20220008565A