Full-level balanced management and control method and system for energy storage battery system
Through a method based on deep reinforcement learning, we construct full-level state space vectors and action space vectors to solve the inconsistency problem of large-scale energy storage battery systems, achieve balanced control of all levels, and improve the safety, reliability and operational efficiency of the system.
Patent Information
- Application Number
- CN202510688385.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, large-scale energy storage battery systems have significant imbalance problems at the battery box, battery cluster, battery stack and energy storage station levels, resulting in reduced system safety, reliability and service life. Existing methods lack comprehensive consideration and coordinated control at all levels.
A method based on deep reinforcement learning is used to construct full-level state space vectors and action space vectors. By building consistency, timeliness and energy consumption reward functions, balanced management and control of each level is achieved. Taking into account the temperature, SOC, SOE, SOH and other status information of the battery system, a dual deep Q network algorithm is used for real-time adjustment.
Effectively solve the inconsistency problem of batteries at all levels, improve the safety and reliability of the battery system, improve energy utilization and charging and discharging efficiency, reduce energy consumption, achieve efficient use of energy, and improve the overall performance and operating benefits of large-scale energy storage battery systems.
Smart Images

Figure CN120638538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large-scale energy storage battery systems, and in particular to a full-level balancing control method and system for large-scale energy storage battery systems based on deep reinforcement learning, specifically covering balancing control at the battery box level, battery cluster level, battery stack level, and energy storage station level. Background Art
[0002] With the rapid development of new energy technologies, large-scale energy storage battery systems are increasingly being used in areas such as power system peak regulation and renewable energy consumption. However, during operation, battery systems are affected by factors such as production processes, operating environments, and aging. This can lead to significant imbalances at all levels, including battery boxes, battery clusters, battery stacks, and energy storage sites. This severely limits the safety, reliability, and service life of these systems.
[0003] At the battery box level, a battery box is composed of multiple battery cells. Inconsistencies between battery cells manifest as differences in temperature (T), state of charge (SOC), state of energy (SOE), and state of health (SOH). This inconsistency can lead to overcharging or over-discharging of some battery cells, shortening their service life and reducing the overall performance and safety of the battery box. At the battery cluster level, multiple battery boxes form a battery cluster. Inconsistencies between battery clusters can lead to uneven charge and discharge currents across the clusters, affecting their charge and discharge efficiency and energy utilization, and ultimately reducing the performance and life of the entire cluster. At the battery stack level, multiple battery stacks within the same energy storage station can also experience inconsistencies due to factors such as inconsistencies within their internal battery clusters and connection methods. This can lead to irrational energy distribution between the stacks and reduce the overall operational efficiency of the energy storage station. At the energy storage station level, within a local power grid, different equipment configurations, operating strategies, and usage across different energy storage stations can lead to inconsistencies between stations, impacting the energy scheduling and stability of the local power grid.
[0004] In the existing technology, most balancing control methods only focus on the inconsistency problem of a single level, lack comprehensive consideration and coordinated control of all levels, and cannot effectively solve the inconsistency problem of all levels of large-scale energy storage battery systems. Summary of the Invention
[0005] To address the deficiencies in the prior art, the present invention provides a full-level balancing control method and system for a large-scale energy storage battery system based on deep reinforcement learning. This method can comprehensively consider the status information of all levels of the battery system, adjust the balancing strategy in real time, and effectively solve the inconsistency problems at each level. This method improves the overall performance and operating efficiency of the large-scale energy storage battery system.
[0006] The present invention adopts the following technical solutions.
[0007] The present invention proposes a full-level balancing control method for an energy storage battery system. The various levels of the energy storage battery system include: battery box level, battery cluster level, battery stack level, and energy storage station level, including:
[0008] Obtain the state vectors of each level to form the state space vector of the entire energy storage battery system;
[0009] Constructing action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system to form the action space vectors of the entire level of the energy storage battery system;
[0010] Using the state variables of each layer, we construct the reward function of each layer and the reward function of the entire energy storage battery system.
[0011] A dual-depth Q network algorithm is used to determine the optimal action space vector based on the current state space vector, and to perform full-level balancing control of the energy storage battery system.
[0012] At the battery box level, the state variables of each battery box are collected to construct the state vector of each battery box level; the state variables of the battery box include voltage, temperature, SOC, SOE, SOH, allowed charging power and allowed discharging power;
[0013] Based on the state vector of each battery box level, determine the state variables of the battery cluster, including voltage, maximum and minimum cell temperature, SOC, SOE, SOH, allowable charge power, and allowable discharge power. Utilize the determined state variables of the battery cluster and the parallel status and current of the battery cluster to construct the state vector of each battery cluster level.
[0014] Based on the state vector of each battery cluster level, the state variables of the battery stack are determined, including the maximum and minimum cell temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power; the state vector of each battery stack level is constructed using the determined state variables of the battery stack and the grid connection status, bus voltage, and current after grid connection.
[0015] Based on the state vector of each battery stack level, determine the state variables of the energy storage station, including SOC, SOE, SOH, allowable charging power, and allowable discharging power; use the determined state variables of the energy storage station and the grid connection status of the energy storage station to construct the state vector of each energy storage station level;
[0016] The state space vector of the entire energy storage battery system is constructed using the state vectors of all battery box levels, all battery cluster levels, all battery stack levels, and all energy storage station levels.
[0017] Based on the state vector of each battery cluster level, the state variables of the battery stack are determined to include:
[0018]
[0019] Where, are the voltage, maximum cell temperature, and minimum cell temperature of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; are the voltage and temperature of the qth battery box in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid; q = 1, 2, …, np, where n is the number of battery boxes in each battery cluster and p is the number of batteries in each battery box; k = 1, 2, …, m, where m is the number of battery clusters in each battery stack; j = 1, 2, …, g, where g is the number of battery stacks in the energy storage station; i = 1, 2, …, h, where h is the number of energy storage stations in the local power grid;
[0020]
[0021] Where, is the SOC of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOC m is the lower limit of the SOC reference value, SOC M is the upper limit of the SOC reference value; is the SOC of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0022]
[0023] Where, is the SOE of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOE m is the lower limit of SOE reference value, SOE M is the upper limit of the SOE reference value; is the SOE of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0024]
[0025] Where, is the SOH of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; is the SOH of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0026]
[0027] Where, are the allowed charging power and discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; They are respectively the allowed charging power and the allowed discharging power of the qth battery box in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid.
[0028] Based on the state vector of each battery cluster level, the state variables of the battery stack are determined to include:
[0029]
[0030] Where, are the maximum and minimum cell temperatures of the j-th battery stack in the i-th energy storage station in the local power grid; is the temperature of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0031]
[0032] Where, is the SOC of the j-th battery stack in the i-th energy storage station in the local power grid;
[0033]
[0034] Where, is the SOE of the j-th battery stack in the i-th energy storage station in the local power grid;
[0035]
[0036] Where, is the SOH of the j-th battery stack in the i-th energy storage station in the local power grid;
[0037]
[0038] Where, are the allowed charging power and discharging power of the j-th battery stack in the i-th energy storage station in the local power grid respectively; They are respectively the allowed charging power and the allowed discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid.
[0039] Based on the state vector of each battery stack level, the state variables of the energy storage station are determined to include:
[0040]
[0041]
[0042] Where, are the grid connection status, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the i-th energy storage station in the local power grid, is the rated charge capacity of the jth battery stack in the i-th energy storage station in the local power grid; is the rated energy capacity of the jth battery stack in the i-th energy storage station in the local power grid.
[0043] At the battery box level, the balancing state of each battery in the battery box is used as the action parameter to construct the action space vector of each battery box level; wherein, the action space vector is (0) or (1) in passive balancing, and the action space vector is (0) or (1) or (-1) in active balancing, where (0) means no action is executed, (1) means discharge action is executed, and (-1) means charge action is executed;
[0044] At the battery cluster level, the action space vectors of each battery cluster level are constructed using the battery cluster parallel operation and charge / discharge power as action parameters. The battery cluster parallel operation is set to 0 to indicate that the battery cluster is disconnected from the battery stack voltage bus, and 1 to indicate that the battery cluster is connected to the battery stack voltage bus. The charge / discharge power is the charge power or the discharge power.
[0045] At the battery stack level, the action space vectors of each battery stack level are constructed using the battery stack grid connection action and charge and discharge power as action parameters. The battery stack grid connection action is set to 0 to indicate that the battery stack is disconnected from the energy storage station grid, and 1 to indicate that the battery stack is connected to the energy storage station grid.
[0046] At the energy storage station level, the grid-connected action and charge / discharge power of the energy storage station are used as action parameters to construct the action space vectors of each energy storage station level. The grid-connected action of the energy storage station is set to 0, indicating that the energy storage station is disconnected from the local power grid, and 1, indicating that the energy storage station is connected to the local power grid.
[0047] The action space vectors of all levels of the energy storage battery system are constructed by using the action space vectors of all battery box levels, the action space vectors of all battery cluster levels, the action space vectors of all battery stack levels, and the action space vectors of all energy storage station levels.
[0048] Using the state variables of each layer, the consistency reward function of each layer is constructed separately; the sum of the consistency reward functions of each layer is used as the consistency reward function of the entire layer of the energy storage battery system;
[0049] Using the state variables of each level, the timeliness reward function of each level is constructed separately; the sum of the timeliness reward functions of each level is used as the timeliness reward function of the entire energy storage battery system;
[0050] Using the state variables of each level, the energy consumption reward function of each level is constructed separately; the sum of the energy consumption reward functions of each level is used as the energy consumption reward function of the entire energy storage battery system;
[0051] The sum of the consistency reward function, timeliness reward function and energy consumption reward function of the energy storage battery system at all levels is used as the reward function of the energy storage battery system at all levels.
[0052] The consistency reward function at the battery box level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery cells in the battery box;
[0053] The consistency reward function at the battery cluster level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery clusters within the battery stack;
[0054] The consistency reward function at the battery stack level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery stack within the energy storage station;
[0055] The consistent reward function at the energy storage site level includes the SOC dispersion, SOE dispersion, and SOH dispersion of the energy storage site within the local power grid;
[0056] The root mean square deviation of each parameter is used to express the degree of dispersion.
[0057] The timeliness reward function at the battery box level includes the degree of change of the temperature of the battery cells in the battery box over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time;
[0058] The time-effectiveness reward function at the battery cluster level includes the degree of change of the temperature of the battery cluster in the battery stack over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time;
[0059] The timeliness reward function at the battery stack level includes the degree of change of the battery stack temperature, SOC, SOE and SOH over time.
[0060] The timeliness reward function at the energy storage station level includes the degree of change of the SOC, SOE, and SOH of the energy storage station in the local power grid over time.
[0061] The energy consumption reward functions are at the battery box level, battery cluster level, battery stack level, and energy storage station level respectively;
[0062]
[0063] Where, It is the energy consumption per unit calculation time under the battery box level balancing action, is the balancing current at the battery box level, is the passive balancing resistor at the battery box level, are the voltage and charge / discharge efficiency of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, Δt is the unit calculation time, and π is a weight coefficient greater than 0. is the action space vector of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0064]
[0065] Where, is the energy consumption per unit calculation time under the battery cluster level balancing action, is the circulating current of the battery cluster, are the voltage, charge and discharge efficiency, and charge and discharge power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, is a weight coefficient greater than 0, is the parallel operation status of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0066]
[0067] In the formula, θ is a weight coefficient greater than 0, are the grid connection status, charge and discharge efficiency, and charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid;
[0068]
[0069] Where, is a weight coefficient greater than 0, P i 站 are the grid connection status, charging and discharging efficiency, and charging and discharging power of the i-th energy storage station in the local power grid.
[0070] The present invention also proposes a full-level balancing control system for an energy storage battery system. The various levels of the energy storage battery system include: battery box level, battery cluster level, battery stack level, and energy storage station level, including:
[0071] The state space construction module is used to obtain the state vectors of each level to form the state space vectors of the entire energy storage battery system;
[0072] An action space construction module is used to construct action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system, so as to form an action space vector of the entire level of the energy storage battery system;
[0073] The reward function establishment module is used to use the state variables of each layer to construct the reward function of each layer and the reward function of the entire energy storage battery system;
[0074] The balancing control module is used to use a dual-depth Q network algorithm to determine the optimal action space vector based on the current state space vector and perform full-level balancing control of the energy storage battery system.
[0075] The beneficial effects of the present invention are that, compared with the prior art, at least the method and system for full-level balancing control of a large-scale energy storage battery system based on deep reinforcement learning proposed in the present invention comprehensively considers the state information of the battery system at the battery box level, battery cluster level, battery stack level, and energy storage station level. By constructing state space vectors and action space vectors for the full level, and constructing reward functions based on the three dimensions of consistency, timeliness, and energy consumption, it achieves real-time adjustment of the balancing control strategy at each level. In terms of reliability, the present invention considers the problems of uneven charge and discharge currents between battery clusters, unreasonable energy distribution between battery stacks, and inconsistency between energy storage stations, and implements coordinated control to improve the overall reliability of the battery system. In terms of improving operational efficiency, by constructing a reward function to optimize the decision-making of intelligent agents, taking into account consistency, timeliness, and energy consumption, the energy utilization rate and charge and discharge efficiency of the battery system are improved, energy consumption is reduced, and efficient energy utilization is achieved, thereby improving the overall performance and operational efficiency of the large-scale energy storage battery system, enabling it to operate more stably and efficiently in areas such as power system peak regulation and renewable energy consumption.
[0076] There are significant differences in technical solutions and actual application effects, resulting in a series of outstanding technical effects.
[0077] Differences: Most existing technologies only focus on the inconsistency problem of a single layer, lacking comprehensive consideration and coordinated control of all layers. The present invention adopts a deep reinforcement learning algorithm to comprehensively consider the status information of all layers of the battery system, including the battery box level, battery cluster level, battery stack level, and energy storage station level, such as the temperature, SOC, SOE, SOH, etc. of batteries at each layer. By constructing the state space vector and action space vector of the entire layer and constructing the reward function from the three dimensions of consistency, timeliness, and energy consumption, real-time adjustment of the balance control strategy of each layer is achieved.
[0078] Technical effect: From the perspective of safety, it effectively solves the inconsistency problem of batteries at all levels, avoids overcharging or over-discharging of battery cells, extends the service life of battery cells, and improves the safety of the entire battery system. For example, at the battery box level, by real-time monitoring and adjusting the battery status, the safety risks caused by inconsistent cells are reduced. In terms of reliability, the present invention takes into account the problems of unbalanced charging and discharging currents between battery clusters, unreasonable energy distribution between battery stacks, and inconsistency between energy storage sites, and performs collaborative control to improve the overall reliability of the battery system. In terms of improving operational efficiency, by constructing a reward function to optimize the decision-making of intelligent agents, taking into account consistency, timeliness and energy consumption, the energy utilization rate and charging and discharging efficiency of the battery system are improved, energy consumption is reduced, and efficient energy utilization is achieved, thereby improving the overall performance and operational efficiency of large-scale energy storage battery systems, so that they can operate more stably and efficiently in the fields of power system peak regulation and renewable energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 This is a schematic diagram of the battery box balancing action in an embodiment of the present invention;
[0080] Figure 2 This is a schematic diagram of a battery cluster balancing operation according to an embodiment of the present invention;
[0081] Figure 3 2. It is a schematic diagram of the battery stack balancing action in an embodiment of the present invention;
[0082] Figure 4 This is a schematic diagram of the balancing action of the energy storage station in an embodiment of the present invention;
[0083] Figure 5 This is a schematic diagram of the framework of the full-level balancing control method for large-scale energy storage battery systems based on deep reinforcement learning proposed in the present invention. DETAILED DESCRIPTION
[0084] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only part of the embodiments of the present invention, not all of them. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0085] The present invention provides a full-level balancing control method for a large-scale energy storage battery system based on deep reinforcement learning, which is applied to all levels of the energy storage system, including the battery box level, battery cluster level, battery stack level, and energy storage station level.
[0086] like Figure 1 As shown, the method includes:
[0087] Step 1: Obtain the state variables of each level of the energy storage system and construct the state vectors of each level of the energy storage system respectively; use the state vectors of each level of the energy storage system to construct the state space vector of the entire level of the energy storage battery system.
[0088] This invention improves upon existing technologies in constructing state vectors at each level. Existing technologies often focus solely on inconsistencies at a single level, lacking comprehensive consideration and coordinated control across all levels. This makes it difficult to fully reflect the overall operating status of large-scale energy storage battery systems when constructing state vectors. This invention, however, considers multiple state variables at each level, starting from the battery box, battery cluster, battery stack, and energy storage station levels, to construct state vectors.
[0089] Battery box level: The state vector is constructed by collecting state variables such as temperature, SOC, SOE, SOH, allowable charging power, and allowable discharge power of each battery. Compared with existing technologies, this more comprehensively covers the key operating parameters of the battery, accurately reflects the operating status of each battery in the battery box, and provides an accurate basis for subsequent balancing control of individual batteries.
[0090] At the battery cluster level, this factor not only considers common variables like voltage and SOC, but also incorporates the maximum and minimum cell temperatures, the cluster's parallel state, and current. This allows the constructed state vector to more accurately describe the operating state of the battery cluster, particularly considering the impact of parallel state on cluster operation. This facilitates more rational cluster-level balancing control and avoids control errors caused by insufficient state vector information.
[0091] At the battery stack and energy storage site levels, a state vector is similarly constructed by integrating multiple state variables, such as the battery stack's grid-connected status, bus voltage and current, and the energy storage site's grid-connected status. This multivariable fusion approach reflects the overall operational status of the energy storage system at a higher level, laying the foundation for coordinated control at all levels.
[0092] Benefits: Comprehensively and accurately reflect the operating conditions of each level of the energy storage system, provide rich and accurate data for deep reinforcement learning algorithms, enable intelligent agents to make decisions based on more complete information, and achieve more effective full-level balanced management and control. By comprehensively considering the state variables of each level to construct a state vector, the battery system can be analyzed and controlled as an organic whole, solving the problem of lack of comprehensive consideration and coordinated control at all levels in existing technologies, improving the overall performance and operating efficiency of large-scale energy storage battery systems, enhancing the safety and reliability of the system, and extending its service life. The steps of the deep reinforcement learning algorithm are as follows: First, initialize the environment of each level of the energy storage system, clarify the state, action space and reward function;
[0093] Specifically, step 1 includes:
[0094] Step 1.1: At the battery box level, collect the state variables of each battery box and construct the state vector of each battery box level; the state variables of the battery box include voltage, temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power;
[0095] At the battery box level, the voltage, temperature, SOC, SOE, SOH, and SOP (including the allowed charging power and allowed discharging power) of each battery box are collected and used as state variables to describe the operating status of the battery box. The state variables of each battery box are used to construct the state vector of each battery box level:
[0096]
[0097] Where, is the state vector of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, are the voltage, temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the qth battery box in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid, respectively. q = 1, 2, …, np, where n is the number of battery boxes in each battery cluster and p is the number of batteries in each battery box.
[0098] Step 1.2: Based on the state vectors of each battery box, determine the state variables of the battery cluster. The determined state variables of the battery cluster include voltage, maximum cell temperature, minimum cell temperature, SOC, SOE, SOH, allowable charge power, and allowable discharge power. The state vectors of each battery cluster are constructed using the determined state variables of the battery cluster and the parallel state and current of the battery cluster.
[0099] Determining the state variables of the battery stack based on the state vectors of each battery cluster level includes:
[0100]
[0101] Where, are the voltage, maximum cell temperature, and minimum cell temperature of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0102]
[0103] Where, is the SOC of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOC m The lower limit of the SOC reference value is 0% in the embodiment. Mis the upper limit of the SOC reference value, which is 100% in the embodiment;
[0104]
[0105] Where, is the SOE of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOE m The lower limit of SOE reference value is 0% in the embodiment. M The upper limit of the SOE reference value is 100% in the embodiment;
[0106]
[0107] Where, is the SOH of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0108]
[0109] Where, are the allowed charging power and discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, n is the number of battery boxes in each battery cluster, and p is the number of batteries in each battery box;
[0110] To obtain the maximum value function, is the minimum value function.
[0111] At the battery cluster level, the parallel status, voltage, current, maximum battery temperature, minimum battery temperature, cluster SOC, cluster SOE, cluster SOH, and cluster SOP (including allowable charging power and allowable discharging power) of each battery cluster are collected and used as state variables to describe the operating status of the battery cluster. Using the above states as the state parameters of each battery cluster, the state vector of each battery cluster level is constructed:
[0112]
[0113] Where, is the state vector of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, are the parallel state, voltage, current, maximum cell temperature, minimum cell temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, respectively. The value 0 indicates that the battery cluster is disconnected from the battery stack voltage bus, and the value 1 indicates that the battery cluster is connected to the battery stack voltage bus. k = 1, 2, ..., m, where m is the number of battery clusters in each battery stack.
[0114] Step 1.3: Determine the state variables of the battery stack based on the state vector of each battery cluster level. The determined battery stack state variables include the maximum and minimum cell temperatures, state of charge (SOC), state of exhaustion (SOE), state of hydration (SOH), allowable charge power, and allowable discharge power. Construct the state vector of each battery stack level using the determined battery stack state variables and the grid connection status, post-grid bus voltage, and current.
[0115]
[0116] Where, are the maximum and minimum cell temperatures of the j-th battery stack in the i-th energy storage station in the local power grid;
[0117]
[0118] Where, is the SOC of the j-th battery stack in the i-th energy storage station in the local power grid;
[0119]
[0120] Where, is the SOE of the j-th battery stack in the i-th energy storage station in the local power grid;
[0121]
[0122] Where, is the SOH of the j-th battery stack in the i-th energy storage station in the local power grid;
[0123]
[0124] Where, are respectively the allowed charging power and the allowed discharging power of the j-th battery stack in the i-th energy storage station in the local power grid.
[0125] At the battery stack level, the grid connection status, total voltage, total current, maximum cell temperature, minimum cell temperature, stack SOC, stack SOE, stack SOH, and stack SOP (including allowable charging power and allowable discharge power) of each battery stack are collected and used as state variables to describe the operating status of the battery stack. The above states are used as the state parameters of each battery stack to construct the state vector of each battery stack level:
[0126]
[0127] Where, is the state vector of the j-th battery stack in the i-th energy storage station in the local power grid, are the grid-connected status, bus voltage, current, maximum cell temperature, minimum cell temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the jth battery stack in the i-th energy storage station in the local power grid, respectively. Taking 0 means that the battery stack leaves the station grid, taking 1 means that the battery stack is connected to the station grid, j = 1, 2, ..., g, g is the number of battery stacks in the energy storage station.
[0128] Step 1.4: Determine the state variables of the energy storage station based on the state vectors of each battery stack level. The determined state variables of the energy storage station include SOC, SOE, SOH, allowable charging power, and allowable discharging power. Utilize the determined state variables of the energy storage station and the grid connection status of the energy storage station to construct the state vectors of each energy storage station level.
[0129]
[0130] Where, are the grid connection status, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the i-th energy storage station in the local power grid, is the rated charge capacity of the jth battery stack in the i-th energy storage station in the local power grid, in Ah; is the rated energy capacity of the jth battery stack in the i-th energy storage station in the local power grid, in kWh;
[0131] At the energy storage station level, the grid connection status, station SOC, station SOE, station SOH, and station SOP (including the allowed charging power and allowed discharging power) are collected and used as state variables to describe the operating status of the energy storage station. The above states are used as the state parameters of each energy storage station to construct the state vector of each energy storage station:
[0132]
[0133] Where, is the state vector of the i-th energy storage station in the local power grid, are the grid connection status, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the i-th energy storage station in the local power grid, respectively. Taking 0 means that the energy storage station leaves the station grid, and taking 1 means that the energy storage station is connected to the station grid. i = 1, 2, …, h, where h is the number of energy storage stations in the local grid.
[0134] In step 1.5, the full-level state space vector of the large-scale energy storage battery system is constructed using the state vectors of all battery box levels, all battery cluster levels, all battery stack levels, and all energy storage station levels, as follows:
[0135]
[0136] Where s is the full-level state space vector of the energy storage battery system, h is the number of energy storage sites in the local power grid, g is the number of battery stacks in the energy storage site, m is the number of battery clusters in the battery stack, n is the number of battery boxes in each battery cluster, and p is the number of battery cells in each battery box.
[0137] Step 2: construct action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system, and use the action space vectors of each level of the energy storage system to form the action space vectors of the entire level of the energy storage battery system.
[0138] In the process of constructing action space vectors at each level, the present invention significantly improves upon existing technologies. This is due to the complexity of large-scale energy storage battery systems and the shortcomings of existing technologies in handling multi-level collaborative control. Existing technologies often only set actions for a single level, lacking a comprehensive consideration of the system as a whole, making it difficult to achieve efficient collaborative control at all levels. However, the present invention fully considers the different characteristics and requirements of the battery box level, battery cluster level, battery stack level, and energy storage station level, constructing a more comprehensive, reasonable, and mutually coordinated action space vector system.
[0139] Battery box level: Using the balancing state of each battery in the battery box as the action parameter, an action space vector is constructed, and the constraints during balancing and the balancing limits of the parallel energy storage array when the battery clusters are not in parallel are clarified. Compared with existing technologies, this action setting that is accurate to the individual battery can accurately control the differences between the batteries in the battery box, effectively avoiding overcharge and over-discharge problems caused by battery inconsistency, extending battery life, and improving the overall performance of the battery box. At the same time, the special limitations of the parallel energy storage array also ensure the stability of the battery cluster when it is in parallel, avoiding the impact of battery box-level balancing on the total voltage of the battery cluster, and thus improving the reliability of the entire energy storage system.
[0140] At the battery cluster level, the action space vector is constructed using the battery cluster parallel operation and charge / discharge power as action parameters. Charge / discharge power constraints are set separately for parallel and string energy storage arrays. Compared to existing technologies, this fully considers the characteristics of different energy storage array structures, making the construction of the action space vector more targeted and enabling flexible regulation based on different energy storage system configurations. This optimizes the battery cluster's charge and discharge processes, improves the cluster's charge / discharge efficiency and energy utilization, and thus enhances the overall operational efficiency of the battery cluster.
[0141] At the battery stack and energy storage site levels, action space vectors are constructed using the battery stack grid-connection action and charge / discharge power, and the energy storage site grid-connection action and charge / discharge power as action parameters, respectively, and corresponding power constraints are set. This construction approach manages and controls the energy storage system at a higher level, enabling reasonable control of grid connection and power regulation for both the battery stack and energy storage site. This ensures stable interaction between the energy storage system and the power grid, improves the energy storage system's energy dispatch capability and stability within the local power grid, and enhances the support provided by large-scale energy storage battery systems to the power grid.
[0142] The coordinated action space vectors at each level constructed in this invention enable full-level coordinated control of large-scale energy storage battery systems. This effectively addresses the lack of comprehensive consideration and coordinated control at all levels in existing technologies, significantly improving the system's overall performance and operational efficiency. Precise action settings and constraints enhance system safety and reliability, reduce the risk of failures caused by imbalances, and extend the lifespan of the energy storage system, providing strong support for the widespread application of large-scale energy storage battery systems in areas such as power system peak regulation and renewable energy consumption.
[0143] In the full-level balancing control method of large-scale energy storage battery systems based on deep reinforcement learning, the construction of the full-level state space vector has been completed, laying the foundation for the intelligent agent to perceive the operating status of the system; but only clarifying the state space is not enough to achieve effective control, and it is necessary to further construct an action space that matches it to determine the set of operations that the intelligent agent can execute at each level; the reasonable construction of the action space is directly related to whether the intelligent agent can achieve the full-level balancing control goal by selecting appropriate actions based on the state space information; therefore, according to the characteristics and requirements of different levels of the energy storage system, constructing corresponding action spaces has become a key link in realizing the technical solution of the present invention, and constructing an action space that matches the full-level state space of the large-scale energy storage battery system as the set of operations performed by the intelligent agent at each level.
[0144] Specifically, step 2 includes:
[0145] Step 2.1: At the battery box level, the equilibrium state of each battery in the battery box is used as the action parameter to construct the action space vector of each battery box level;
[0146]
[0147] Where, is the action space vector of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, (0) means no action is performed, (1) means the discharge action is performed, and (-1) means the charge action is performed;
[0148] The kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid is as follows: Figure 1 As shown, the battery box 1 They are the action space vectors of battery cells 1, 2, ..., p respectively; in battery box 2 They are the action space vectors of battery cells p+1, p+2, ..., 2p respectively; are the action space vectors of battery cells np-p+1, np-p+2, ..., np respectively;
[0149] When active balancing is used at the battery box level, the following constraints must be met:
[0150]
[0151] Where A N The maximum number of channels for supplementary charging using external power supply at the battery box level.
[0152] For parallel energy storage arrays, to avoid battery box level balancing, which would cause changes in the total voltage of the battery cluster and affect parallel operation, balancing is not performed when the battery clusters are not connected in parallel. This is indicated as follows:
[0153]
[0154] Step 2.2: At the battery cluster level, the action space vector of each battery cluster level is constructed using the battery cluster parallel operation and charge and discharge power as action parameters;
[0155]
[0156] Where, is the action space vector of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, is the parallel operation of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, The value 0 indicates that the battery cluster is disconnected from the battery stack voltage busbar, and the value 1 indicates that the battery cluster is connected to the battery stack voltage busbar. is the charge and discharge power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid. For the parallel energy storage array, For string energy storage arrays,
[0157] The jth battery stack in the i-th energy storage station in the local power grid is as follows: Figure 2 As shown, is the action space vector of battery cluster 1, is the action space vector of battery cluster 2, is the action space vector of battery cluster m;
[0158] Step 2.3: At the battery stack level, the action space vector of each battery stack level is constructed using the battery stack grid connection action and charge and discharge power as action parameters;
[0159]
[0160] Where, is the action space vector of the jth battery stack in the i-th energy storage station in the local power grid, is the grid-connected action of the jth battery stack in the i-th energy storage station in the local power grid, The value 0 indicates that the battery stack is disconnected from the energy storage station grid, and the value 1 indicates that the battery stack is connected to the energy storage station grid. is the charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid, satisfying
[0161] The i-th energy storage station in the local power grid is as follows: Figure 3 As shown, is the action space vector of battery stack 1, ..., is the action space vector of battery stack g;
[0162] Step 2.4: At the energy storage station level, the action space vectors of each energy storage station level are constructed using the grid-connected action and charge / discharge power of the energy storage station as action parameters.
[0163]
[0164] P i 站 =P i 调度
[0165] Where, is the action space vector of the i-th energy storage station in the local power grid, is the grid-connected action of the i-th energy storage station in the local power grid, is the dispatching and grid-connecting action of the i-th energy storage station in the local power grid, The value 0 indicates that the energy storage station is disconnected from the local power grid, and the value 1 indicates that the energy storage station is connected to the local power grid. i站 is the charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid, P i 调度 is the dispatching charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid, satisfying
[0166] Local power grid Figure 4 As shown, is the action space vector of energy storage station 1, is the action space vector of the energy storage station h;
[0167] In step 2.5, the full-level action space vector of the large-scale energy storage battery system is constructed using the action space vectors of all battery box levels, all battery cluster levels, all battery stack levels, and all energy storage station levels, as follows:
[0168]
[0169] Where a is the full-level action space vector of the energy storage battery system, h is the number of energy storage sites in the local power grid, g is the number of battery stacks in the energy storage site, m is the number of battery clusters in the battery stack, n is the number of battery boxes in each battery cluster, and p is the number of battery cells in each battery box.
[0170] Step 3: Use the state variables of each level of the energy storage system to construct the reward function of each level of the energy storage system respectively; use the sum of the functions of each level of the energy storage system as the reward function of the entire level of the energy storage battery system; among them, the reward function of each level of the energy storage system and the reward function of the entire level of the energy storage battery system both include a consistency reward function, a timeliness reward function and an energy consumption reward function; the sum of the consistency reward function, the timeliness reward function and the energy consumption reward function of the entire level of the energy storage battery system is used as the reward function of the entire level of the energy storage battery system.
[0171] This invention significantly improves upon existing technologies by constructing reward functions at each level. Existing technologies often focus on single-level balance control and lack comprehensiveness and systematicness in their reward function construction, failing to comprehensively consider the complex operational conditions of all levels of the energy storage system. This invention, however, achieves a breakthrough by constructing reward functions for each level and across all levels, focusing on consistency, timeliness, and energy consumption.
[0172] Comprehensive consideration of dimensions: The reward functions constructed by existing technologies often only focus on a single indicator. For example, they may only focus on the balance of battery power and ignore other important factors. The present invention comprehensively considers the three dimensions of consistency, timeliness and energy consumption, and comprehensively evaluates the effects of the execution of actions at each level. In terms of consistency, consistency reward functions are constructed for the battery box, battery cluster, battery stack and energy storage station levels respectively, considering the average values of temperature, SOC, SOE and SOH at each level, so as to make the battery status at each level more consistent, reduce local overcharge and over-discharge caused by inconsistency, and improve the overall safety and stability of the battery system. In the dimension of timeliness, a reward function is constructed based on the time rate of change of state variables at each level to ensure that the energy storage system can respond to state changes in a timely manner, quickly adjust the operation strategy, enhance the dynamic response capability of the system, and improve the system operation efficiency. Constructing a reward function from the energy consumption dimension effectively controls the energy consumption of balancing actions at each level, optimizes energy utilization efficiency, and reduces operating costs.
[0173] Full-level collaborative construction: Existing technologies lack collaborative consideration of all levels of the energy storage system. The reward functions of each level are independent of each other, and overall optimization cannot be achieved. The present invention constructs a full-level reward function based on the sum of the reward functions of each level, closely linking the battery box level, battery cluster level, battery stack level, and energy storage station level. This construction method enables the intelligent agent to no longer be limited to the optimization of a single level when making decisions, but to achieve full-level collaborative control from the perspective of the entire energy storage system, thereby improving the overall performance and operating efficiency of large-scale energy storage battery systems.
[0174] The present invention constructs a reward function through comprehensive dimensions and full-level collaborative construction, providing the intelligent agent with more comprehensive and accurate feedback signals, guiding the intelligent agent to learn and optimize its decision-making. Under such a reward mechanism, the intelligent agent can gradually explore and master the optimal balance management and control strategy, effectively solving the inconsistency problem of all levels of large-scale energy storage battery systems. This not only improves the safety, reliability and service life of the battery system, but also enhances the adaptability of the energy storage system in the power system, enabling it to better serve application scenarios such as power system peak regulation and renewable energy consumption, providing strong support for the widespread application and efficient operation of large-scale energy storage battery systems.
[0175] After constructing action space vectors at the battery box, cluster, stack, and energy storage station levels, the agent has a clear set of executable operations at each level. However, to guide the agent's learning and optimize its decisions, achieving the core goal of full-level balanced management and control of large-scale energy storage battery systems, a corresponding reward function must be constructed. As a feedback mechanism in the agent's interaction with the environment, its design directly influences the direction and efficiency of the agent's strategy learning. By properly defining the reward function, the effects of actions executed at each level can be quantified, thereby prompting the agent to gradually explore and master the optimal balanced management and control strategy.
[0176] The present invention constructs a full-level reward function from three dimensions: consistency, timeliness, and energy consumption of the large-scale energy storage battery system.
[0177] Specifically, step 3 includes:
[0178] Step 3.1: Use the state variables of each level of the energy storage system to construct the consistency reward function of each level of the energy storage system. The sum of the consistency reward functions of each level of the energy storage system is used as the consistency reward function of the entire level of the energy storage battery system.
[0179] In terms of consistency, the full-level reward function of the large-scale energy storage battery system is as follows:
[0180]
[0181] Where R 一致性 is the consistency reward function of the energy storage battery system, These are the consistency reward functions at the battery box level, battery cluster level, battery stack level, and energy storage station level;
[0182] The consistency reward function at the battery box level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery cells in the battery box; the dispersion is expressed using the root mean square deviation of each parameter;
[0183]
[0184] In the formula, α, β, χ, and δ are all weight coefficients greater than 0. are the average values of temperature, SOC, SOE, and SOH at the battery box level;
[0185] It is the degree of dispersion of the battery cell temperature in the battery box, which is reflected by the root mean square deviation. The minus sign in front indicates that the smaller the temperature dispersion (that is, the better the consistency), the higher the reward value.
[0186] It is the degree of dispersion of the SOC of the battery cells in the battery box, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOC dispersion (that is, the better the consistency), the higher the reward value.
[0187] It is the degree of dispersion of the SOE of the battery cells in the battery box, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of dispersion of SOE (that is, the better the consistency), the higher the reward value.
[0188] It is the degree of dispersion of the SOH of the battery cells in the battery box, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOH dispersion (that is, the better the consistency), the higher the reward value.
[0189] The consistency reward function at the battery cluster level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery clusters within the battery stack;
[0190]
[0191] Where, ε, φ, γ is a weight coefficient greater than 0, are the average values of temperature, SOC, SOE, and SOH at the battery cluster level;
[0192] It is the degree of dispersion of the temperature of the battery cluster units in the battery stack, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the temperature dispersion (that is, the better the consistency), the higher the reward value.
[0193] It is the degree of dispersion of the SOC of the battery cluster units in the battery stack, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOC dispersion (that is, the better the consistency), the higher the reward value.
[0194] It is the degree of dispersion of the SOE of the battery cluster units in the battery stack, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of dispersion of SOE (that is, the better the consistency), the higher the reward value.
[0195] It is the degree of dispersion of the SOH of the battery cluster units in the battery stack, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of SOH dispersion (that is, the better the consistency), the higher the reward value.
[0196] The consistency reward function at the battery stack level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery stack within the energy storage station;
[0197]
[0198] In the formula, η, ι, κ, and λ are all weight coefficients greater than 0. are the average values of temperature, SOC, SOE, and SOH at the battery stack level;
[0199] It is the degree of dispersion of the battery stack unit temperature in the battery station, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the temperature dispersion (that is, the better the consistency), the higher the reward value.
[0200] It is the degree of dispersion of the SOC of the battery stack units in the battery station, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOC dispersion (that is, the better the consistency), the higher the reward value.
[0201] It is the degree of dispersion of the SOE of the battery stack units in the battery station, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of dispersion of SOE (that is, the better the consistency), the higher the reward value.
[0202] It is the degree of dispersion of the SOH of the battery stack units in the battery station, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOH dispersion (that is, the better the consistency), the higher the reward value.
[0203] The consistent reward function at the energy storage site level includes the SOC dispersion, SOE dispersion, and SOH dispersion of the energy storage site within the local power grid;
[0204]
[0205] Wherein, μ, ν, and ο are all weight coefficients greater than 0, They are the average values of SOC, SOE, and SOH at the energy storage site level.
[0206] It is the degree of dispersion of the SOC of the energy storage station in the local power grid, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the SOC dispersion (that is, the better the consistency), the higher the reward value.
[0207] It is the degree of dispersion of the SOE of energy storage stations in the local power grid, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of dispersion of SOE (that is, the better the consistency), the higher the reward value.
[0208] It is the degree of dispersion of the SOH of the energy storage station in the local power grid, which is reflected by the root mean square deviation. The negative sign in front indicates that the smaller the degree of dispersion of SOH (that is, the better the consistency), the higher the reward value.
[0209] Step 3.2: Using the state variables of each level of the energy storage system, construct the timeliness reward function of each level of the energy storage system respectively; and use the sum of the timeliness reward functions of each level of the energy storage system as the timeliness reward function of the entire level of the energy storage battery system;
[0210] In terms of timeliness, the full-level reward function of the large-scale energy storage battery system is as follows:
[0211]
[0212] Where R 时效性 is the timeliness reward function of the energy storage battery system, These are the time-effectiveness reward functions at the battery box level, battery cluster level, battery stack level, and energy storage station level;
[0213] The timeliness reward function at the battery box level includes the degree of change of the temperature of the battery cells in the battery box over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time;
[0214] The time-effectiveness reward function at the battery cluster level includes the degree of change of the temperature of the battery cluster in the battery stack over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time;
[0215] The timeliness reward function at the battery stack level includes the degree of change of the battery stack temperature, SOC, SOE and SOH over time.
[0216] The timeliness reward function at the energy storage station level includes the degree of change of the SOC, SOE, and SOH of the energy storage station in the local power grid over time.
[0217] The root mean square deviation of each parameter is used to express the degree of dispersion.
[0218]
[0219] Where, are all weight coefficients greater than 0, are the time-varying rates of temperature, SOC, SOE, and SOH of the qth battery in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid;
[0220] It is the degree of change of the temperature of the battery cells in the battery box over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0221] It is the degree of change of the SOC of the battery cells in the battery box over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0222] It is the degree of change of the SOE of the battery cells in the battery box over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0223] It is the degree of change of the battery cell SOH in the battery box over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0224]
[0225] Where, are all weight coefficients greater than 0, are the time-varying rates of temperature, SOC, SOE, and SOH of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0226] It is the degree to which the temperature of the battery cluster in the battery stack changes over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0227] It is the degree of change of the SOC of the battery cluster in the battery stack over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0228] It is the degree to which the SOE of the battery cluster in the battery stack changes over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0229] It is the degree to which the SOH of the battery cluster in the battery stack changes over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0230]
[0231] Where, ∞、 are all weight coefficients greater than 0, are the time change rates of the temperature, SOC, SOE, and SOH of the j-th battery stack in the i-th energy storage station in the local power grid;
[0232] It is the degree of change of the battery stack temperature in the battery station over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0233] It is the degree of change of the battery stack SOC in the battery station over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0234] It is the degree of change of the SOE of the battery stack in the battery station over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0235] It is the degree of change of the battery stack SOH in the battery station over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0236]
[0237] Where, are all weight coefficients greater than 0, are the time change rates of temperature, SOC, SOE, and SOH of the i-th energy storage station in the local power grid.
[0238] It is the degree of change of the SOC of the battery station in the local power grid over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0239] It is the degree of change of the SOE of the battery station in the local power grid over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0240] It is the degree of change of the SOH of the battery station in the local power grid over time. The absolute value indicates that the greater the change, the more it is affected by the weight. The positive sign in front indicates that the larger the amplitude (that is, the better the effectiveness), the higher the reward value.
[0241] Step 3.3: Using the state variables of each level of the energy storage system, construct the energy consumption reward function of each level of the energy storage system respectively; and use the sum of the energy consumption reward functions of each level of the energy storage system as the energy consumption reward function of the entire level of the energy storage battery system;
[0242] In terms of energy consumption, the full-level reward function of a large-scale energy storage battery system is as follows:
[0243]
[0244] Where R 能耗 is the energy consumption reward function of the energy storage battery system, The energy consumption reward functions are at the battery box level, battery cluster level, battery stack level, and energy storage station level respectively;
[0245]
[0246] Where, It is the energy consumption per unit calculation time under the battery box level balancing action, is the balancing current at the battery box level, is the passive balancing resistor at the battery box level, are the voltage and charge / discharge efficiency of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, Δt is the unit calculation time, and π is a weight coefficient greater than 0. is the action space vector of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0247]
[0248] Where, is the energy consumption per unit calculation time under the battery cluster level balancing action, is the circulating current of the battery cluster, are the voltage, charge and discharge efficiency, and charge and discharge power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, is a weight coefficient greater than 0, is the parallel operation status of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid;
[0249]
[0250] In the formula, θ is a weight coefficient greater than 0, are the grid connection status, charge and discharge efficiency, and charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid;
[0251]
[0252] Where, is a weight coefficient greater than 0, P i 站are the grid connection status, charging and discharging efficiency, and charging and discharging power of the i-th energy storage station in the local power grid.
[0253] Step 3.4: The sum of the consistency reward function, timeliness reward function, and energy consumption reward function of the energy storage battery system is used as the reward function of the energy storage battery system.
[0254] The full-level reward function of the large-scale energy storage battery system is:
[0255] R=R 一致性 +R 时效性 +R 能耗
[0256] In step 4, a dual-depth Q network algorithm is used to determine the optimal action space vector based on the current state space vector, and to perform full-level balancing control of the energy storage battery system.
[0257] The present invention adopts the Double Deep-Q-Network (DDQN) algorithm for deep reinforcement learning training. DDQN includes an evaluation network and a target network; it is a non-restrictive and preferred choice.
[0258] In the full-level balancing control method for large-scale energy storage battery systems, the dual deep Q network (DDQN) algorithm is used for deep reinforcement learning training. Compared with existing technologies, it has significant improvements and brings many benefits:
[0259] Existing technologies often suffer from training instability and susceptibility to local optima when addressing the complexities of energy storage systems. Traditional single-depth Q-network algorithms are susceptible to overestimation bias when estimating target Q values, leading to unstable learning processes and difficulty converging to the optimal strategy. This makes them ineffective in addressing the complex multi-layer, multi-variable environments found in large-scale energy storage battery systems.
[0260] The DDQN algorithm employed in this paper offers a key improvement over existing techniques by introducing an evaluation network and a target network. The evaluation network predicts the Q-value of an action based on the current state, providing a basis for the agent's decision-making. The target network calculates the target Q-value, with a relatively low update frequency, maintaining relative stability. This dual-network structure effectively decouples the calculation of the target Q-value from the estimation of the current policy, reducing the impact of overestimation bias and enhancing the stability of algorithm training.
[0261] Specifically, during experience replay and network updates, when randomly sampling data from the replay buffer to calculate the target value, the evaluation network is used to select actions, and the target network calculates the target Q value. This approach avoids the problem of accumulated deviations caused by using the same network to select actions and calculate target values in a single deep Q network, allowing the algorithm to more accurately learn the optimal strategy when faced with the complex and changing state space and action space of large-scale energy storage battery systems. For example, when processing state changes and action decisions at various levels of battery boxes, battery clusters, battery stacks, and energy storage stations, the DDQN algorithm can more stably update network parameters, improving the algorithm's convergence speed and accuracy.
[0262] Furthermore, after a certain number of training steps, the evaluation network parameters are copied to the target network, further ensuring the algorithm's stability and convergence. Through this mechanism, the target network can regularly obtain updates from the evaluation network while maintaining relative stability, providing a reliable reference for the evaluation network's training and preventing oscillations or local optimal solutions during training.
[0263] Using the DDQN algorithm for deep reinforcement learning training significantly improves the effectiveness of full-level balancing control in large-scale energy storage battery systems. By enhancing training stability and convergence, the intelligent agent can more efficiently learn the optimal strategy for full-level balancing control and more accurately adjust balancing actions at each level, effectively resolving inconsistencies across all levels of the large-scale energy storage battery system and improving overall system performance and operational efficiency.
[0264] Specifically, if Figure 5 As shown, step 4 includes:
[0265] Step 4.1: Initialize the state space vectors, action space vectors, and various reward functions at each level.
[0266] Initialize the energy storage system environment at each level, clarify the state space vector s, action space vector a and reward function R. The weight coefficients α, β, χ, δ, ε, φ, γ, η, ι, κ, λ, μ, ν, ο, ∞、 π, θ、 As model hyperparameters, initialize the agent, build the policy with a dual-depth Q network, and initialize the playback buffer.
[0267] Step 4.2, sampling and interaction;
[0268] The agent is based on the current state space vector s 当前 Select the action space vector a 当前 , based on the selected action space vector a 当前Feedback the new state space vector s 新 And the reward function R corresponding to the new state space vector 新 , will (s 当前 a 当前 s 新 R 新 ) into the playback buffer.
[0269] Step 4.3, experience replay and network update;
[0270] Randomly sample a batch of data from the playback buffer to calculate the target value as follows:
[0271]
[0272] Where Y is the target value, is the discount factor, Q eval To evaluate the network, Q target is the target network, and a is a randomly sampled action space vector.
[0273] After N random samplings, the error between the target value and the predicted value of the evaluation network is calculated as follows:
[0274]
[0275] Where Loss is the error between the target value and the predicted value of the evaluation network, Y n is the target value of the nth random sampling, N is the number of random sampling, s n is the state space vector of the nth random sampling, a n is the action space vector randomly sampled for the nth time;
[0276] Use the back-propagation algorithm to update the evaluation network Q eval parameters, so that the Q value output by the evaluation network is closer to the target value; the present invention updates the evaluation network parameters by randomly sampling data from the playback buffer and using the calculation results to achieve the effect of using the action space vector with a certain degree of randomness to influence the state space vector of the entire system.
[0277] The action space vector is evaluated by the feedback of the state space vector of the entire system. If the balancing effect is good, positive feedback is given; if the balancing effect is poor, negative feedback is given, and the action space vector is adjusted. The evaluation of the action space vector is completed by the target network, and the adjustment of the action space vector is completed by the evaluation network.
[0278] Step 4.4, target network update;
[0279] After M steps of training, the network Q will be evaluated eval The parameters are copied to the target network Q target, synchronizing the parameters of the target network with the evaluation network to ensure algorithm stability and convergence. In this embodiment, M = 500. After a certain number of evaluations by the evaluation network, the target network parameters synchronize with the evaluation network parameters, completing one algorithm iteration. The algorithm continues running, continuously outputting action space vectors, using feedback state to adjust algorithm parameters, and then outputting action space vectors again, repeating the cycle.
[0280] Step 4.5, loop iteration;
[0281] Repeating the aforementioned process of interaction, storage, experience replay, network update, and target network update, the grid search algorithm is used to update model hyperparameters, achieving iterative optimization. The overall algorithm is an iterative calculation. At the beginning of the algorithm, these hyperparameters are assigned initial values. Subsequent algorithm steps then evaluate the action space. This constitutes a set of hyperparameter trials. Then, repeating step 4.5, the hyperparameter combination is changed, and a new round of algorithm iteration is performed to reevaluate the action space, completing the next round of the algorithm.
[0282] The algorithm iterates over and over again, gradually optimizing the hyperparameter combination to improve the consistency, timeliness, and energy efficiency of all-level balancing control.
[0283] In the full-level balancing control method for large-scale energy storage battery systems based on deep reinforcement learning, the evaluation network and target network are key components of the dual deep Q-network (DDQN) algorithm. The two work closely together to promote the stable operation and effective learning of the algorithm.
[0284] The evaluation network predicts the Q-value of an action based on the current state, providing a basis for the agent to select an action in the current state. The target network is responsible for calculating the target Q-value to update the evaluation network's parameters. When calculating the target value, the evaluation network selects an action, and the target network calculates the target Q-value for the corresponding action. These two functions work together to complete the Q-value estimation and update.
[0285] The evaluation network is frequently updated during training, and its parameters are continuously adjusted through the back-propagation algorithm based on the data replayed from experience and the calculated errors to more accurately predict the Q value. The target network is updated less frequently, and the parameters of the evaluation network are usually copied after a certain number of steps (such as 500 steps in the embodiment) of training. This asynchronous update method allows the target network to remain relatively stable, providing a stable and reliable reference for the training of the evaluation network, and avoiding the problem of training instability caused by frequent changes in the target value.
[0286] The evaluation network continuously learns and updates, gradually improving its ability to estimate the value of actions in the current state. Its updated results are regularly transmitted to the target network, enabling it to obtain the latest policy information. Based on this stable information, the target network calculates the target Q value, providing a more reasonable target for the next round of updates of the evaluation network and promoting further optimization of the evaluation network. The two mutually enhance each other and jointly improve the performance of the algorithm.
[0287] The role of the action space output by the algorithm:
[0288] (1) Achieve balanced management and control
[0289] The action space vector is composed of action space vectors at each level, including action decision information at the battery box level, battery cluster level, battery stack level, and energy storage station level. Using these action space vectors, the intelligent agent can perform corresponding balancing actions based on the real-time status of each level of the energy storage system, such as charging and discharging actions at the battery box level, paralleling actions at the battery cluster level, and charge and discharge power adjustment. This achieves balanced management and control of all levels of large-scale energy storage battery systems and resolves inconsistencies at each level.
[0290] (2) Guide system optimization
[0291] The action space provides the agent with a set of actions that can be performed in different states. The algorithm continuously explores and optimizes the action space, guiding the agent to learn the optimal equilibrium control strategy. In this process, the agent continuously adjusts its action selection based on rewards from environmental feedback, enabling the energy storage system to achieve a more optimal operating state in terms of safety, reliability, and energy consumption, thereby improving the overall performance and operational efficiency of large-scale energy storage battery systems.
[0292] (3) Support system decision-making
[0293] This provides a direct basis for decision-making at all levels of the energy storage system. When the system is in different operating states, the algorithm selects appropriate actions from the action space based on the state space vector and learned strategies, achieving precise control of system operation. At the battery stack level, based on information such as the battery stack's grid connection status and SOC, it determines from the action space whether to execute the grid connection action and adjusts the charge and discharge power, ensuring stable interaction between the battery stack and the grid and rational energy distribution.
[0294] The above is only part of the implementation principle of the present invention and does not limit the present invention in any form. Using deep reinforcement learning methods, by establishing state space vectors of state parameters such as temperature, SOC, SOE, SOH, SOP, etc. of large-scale energy storage battery system boxes, clusters, piles, and stations, and constructing full-level balanced action space vectors for boxes, clusters, piles, and stations, a method for achieving full-level balanced management and control falls within the scope of protection of the technical solution of the present invention.
[0295] The present invention also proposes a full-level balancing control system for an energy storage battery system. The various levels of the energy storage battery system include: battery box level, battery cluster level, battery stack level, and energy storage station level, including:
[0296] The state space construction module is used to obtain the state vectors of each level to form the state space vectors of the entire energy storage battery system;
[0297] An action space construction module is used to construct action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system, so as to form an action space vector of the entire level of the energy storage battery system;
[0298] The reward function establishment module is used to use the state variables of each layer to construct the reward function of each layer and the reward function of the entire energy storage battery system;
[0299] The balancing control module is used to use a dual-depth Q network algorithm to determine the optimal action space vector based on the current state space vector and perform full-level balancing control of the energy storage battery system.
[0300] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0301] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0302] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0303] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0304] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. Energy storage battery system full-level balancing control method. The various levels of the energy storage battery system include: The battery box level, battery cluster level, battery stack level, and energy storage station level are characterized by including: Obtain the state vectors of each level to form the state space vector of the entire energy storage battery system; Constructing action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system to form the action space vectors of the entire level of the energy storage battery system; Using the state variables of each layer, we construct the reward function of each layer and the reward function of the entire energy storage battery system. A dual-depth Q network algorithm is used to determine the optimal action space vector based on the current state space vector, and to perform full-level balancing control of the energy storage battery system.
2. The energy storage battery system full-level balancing control method according to claim 1, characterized in that: At the battery box level, the state variables of each battery box are collected to construct the state vector of each battery box level; the state variables of the battery box include voltage, temperature, SOC, SOE, SOH, allowed charging power and allowed discharging power; Based on the state vector of each battery box level, determine the state variables of the battery cluster, including voltage, maximum and minimum cell temperature, SOC, SOE, SOH, allowable charge power, and allowable discharge power. Utilize the determined state variables of the battery cluster and the parallel status and current of the battery cluster to construct the state vector of each battery cluster level. Based on the state vector of each battery cluster level, the state variables of the battery stack are determined, including the maximum and minimum cell temperature, SOC, SOE, SOH, allowed charging power, and allowed discharging power; the state vector of each battery stack level is constructed using the determined state variables of the battery stack and the grid connection status, bus voltage, and current after grid connection. Based on the state vector of each battery stack level, determine the state variables of the energy storage station, including SOC, SOE, SOH, allowable charging power, and allowable discharging power; use the determined state variables of the energy storage station and the grid connection status of the energy storage station to construct the state vector of each energy storage station level; The state space vector of the entire energy storage battery system is constructed using the state vectors of all battery box levels, all battery cluster levels, all battery stack levels, and all energy storage station levels.
3. The energy storage battery system full-level balancing control method according to claim 2, characterized in that: Based on the state vector of each battery cluster level, the state variables of the battery stack are determined to include: Where, are the voltage, maximum cell temperature, and minimum cell temperature of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; are the voltage and temperature of the qth battery box in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid; q = 1, 2, …, np, where n is the number of battery boxes in each battery cluster and p is the number of batteries in each battery box; k = 1, 2, …, m, where m is the number of battery clusters in each battery stack; j = 1, 2, …, g, where g is the number of battery stacks in the energy storage station; i = 1, 2, …, h, where h is the number of energy storage stations in the local power grid; Where, is the SOC of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOC m is the lower limit of the SOC reference value, SOC M is the upper limit of the SOC reference value; is the SOC of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; Where, is the SOE of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, SOE m is the lower limit of SOE reference value, SOE M is the upper limit of the SOE reference value; is the SOE of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; Where, is the SOH of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; is the SOH of the qth battery box in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; Where, are the allowed charging power and discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; They are respectively the allowed charging power and the allowed discharging power of the qth battery box in the kth battery cluster in the jth battery stack in the ith energy storage station in the local power grid.
4. The energy storage battery system full-level balancing control method according to claim 3, characterized in that: Based on the state vector of each battery cluster level, the state variables of the battery stack are determined to include: Where, are the maximum and minimum cell temperatures of the j-th battery stack in the i-th energy storage station in the local power grid; is the temperature of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; Where, is the SOC of the j-th battery stack in the i-th energy storage station in the local power grid; Where, is the SOE of the j-th battery stack in the i-th energy storage station in the local power grid; Where, is the SOH of the j-th battery stack in the i-th energy storage station in the local power grid; Where, are the allowed charging power and discharging power of the j-th battery stack in the i-th energy storage station in the local power grid respectively; They are respectively the allowed charging power and the allowed discharging power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid.
5. The energy storage battery system full-level balancing control method according to claim 4, characterized in that: Based on the state vector of each battery stack level, the state variables of the energy storage station are determined to include: Where, are the grid connection status, SOC, SOE, SOH, allowed charging power, and allowed discharging power of the i-th energy storage station in the local power grid, is the rated charge capacity of the jth battery stack in the i-th energy storage station in the local power grid; is the rated energy capacity of the jth battery stack in the i-th energy storage station in the local power grid.
6. The energy storage battery system full-level balancing control method according to claim 1, characterized in that: At the battery box level, the balancing state of each battery in the battery box is used as the action parameter to construct the action space vector of each battery box level; wherein, the action space vector is (0) or (1) in passive balancing, and the action space vector is (0) or (1) or (-1) in active balancing, where (0) means no action is executed, (1) means discharge action is executed, and (-1) means charge action is executed; At the battery cluster level, the action space vectors of each battery cluster level are constructed using the battery cluster parallel operation and charge / discharge power as action parameters. The battery cluster parallel operation is set to 0 to indicate that the battery cluster is disconnected from the battery stack voltage bus, and 1 to indicate that the battery cluster is connected to the battery stack voltage bus. The charge / discharge power is the charge power or the discharge power. At the battery stack level, the action space vectors of each battery stack level are constructed using the battery stack grid connection action and charge and discharge power as action parameters. The battery stack grid connection action is set to 0 to indicate that the battery stack is disconnected from the energy storage station grid, and 1 to indicate that the battery stack is connected to the energy storage station grid. At the energy storage station level, the grid-connected action and charge / discharge power of the energy storage station are used as action parameters to construct the action space vectors of each energy storage station level. The grid-connected action of the energy storage station is set to 0, indicating that the energy storage station is disconnected from the local power grid, and 1, indicating that the energy storage station is connected to the local power grid. The action space vectors of all levels of the energy storage battery system are constructed by using the action space vectors of all battery box levels, the action space vectors of all battery cluster levels, the action space vectors of all battery stack levels, and the action space vectors of all energy storage station levels.
7. The energy storage battery system full-level balancing control method according to claim 5, characterized in that: Using the state variables of each layer, the consistency reward function of each layer is constructed separately; the sum of the consistency reward functions of each layer is used as the consistency reward function of the entire layer of the energy storage battery system; Using the state variables of each level, the timeliness reward function of each level is constructed separately; the sum of the timeliness reward functions of each level is used as the timeliness reward function of the entire energy storage battery system; Using the state variables of each level, the energy consumption reward function of each level is constructed separately; the sum of the energy consumption reward functions of each level is used as the energy consumption reward function of the entire energy storage battery system; The sum of the consistency reward function, timeliness reward function and energy consumption reward function of the energy storage battery system at all levels is used as the reward function of the energy storage battery system at all levels.
8. The energy storage battery system full-level balancing control method according to claim 7, characterized in that: The consistency reward function at the battery box level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery cells in the battery box; The consistency reward function at the battery cluster level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery clusters within the battery stack; The consistency reward function at the battery stack level includes the temperature dispersion, SOC dispersion, SOE dispersion, and SOH dispersion of the battery stack within the energy storage station; The consistent reward function at the energy storage site level includes the SOC dispersion, SOE dispersion, and SOH dispersion of the energy storage site within the local power grid; The root mean square deviation of each parameter is used to express the degree of dispersion.
9. The energy storage battery system full-level balancing control method according to claim 8, characterized in that: The timeliness reward function at the battery box level includes the degree of change of the temperature of the battery cells in the battery box over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time; The time-effectiveness reward function at the battery cluster level includes the degree of change of the temperature of the battery cluster in the battery stack over time, the degree of change of the SOC over time, the degree of change of the SOE over time, and the degree of change of the SOH over time; The timeliness reward function at the battery stack level includes the degree of change of the battery stack temperature, SOC, SOE and SOH over time. The timeliness reward function at the energy storage station level includes the degree of change of the SOC, SOE, and SOH of the energy storage station in the local power grid over time.
10. The energy storage battery system full-level balancing control method according to claim 9, characterized in that: The energy consumption reward functions are at the battery box level, battery cluster level, battery stack level, and energy storage station level respectively; Where, It is the energy consumption per unit calculation time under the battery box level balancing action, is the balancing current at the battery box level, is the passive balancing resistor at the battery box level, are the voltage and charge / discharge efficiency of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, Δt is the unit calculation time, and π is a weight coefficient greater than 0. is the action space vector of the qth battery in the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; Where, is the energy consumption per unit calculation time under the battery cluster level balancing action, is the circulating current of the battery cluster, are the voltage, charge and discharge efficiency, and charge and discharge power of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid, is a weight coefficient greater than 0, is the parallel operation status of the kth battery cluster in the jth battery stack in the i-th energy storage station in the local power grid; In the formula, θ is a weight coefficient greater than 0, are the grid connection status, charge and discharge efficiency, and charge and discharge power of the jth battery stack in the i-th energy storage station in the local power grid; Where, is a weight coefficient greater than 0, P i 站 are the grid connection status, charging and discharging efficiency, and charging and discharging power of the i-th energy storage station in the local power grid.
11. A full-level balancing control system for an energy storage battery system, wherein each level of the energy storage battery system includes: The battery box level, battery cluster level, battery stack level, and energy storage station level are characterized by including: The state space construction module is used to obtain the state vectors of each level to form the state space vectors of the entire energy storage battery system; An action space construction module is used to construct action space vectors of each level of the energy storage system that match the state vectors of each level of the energy storage system, so as to form an action space vector of the entire level of the energy storage battery system; The reward function establishment module is used to use the state variables of each layer to construct the reward function of each layer and the reward function of the entire energy storage battery system; The balancing control module is used to use a dual-depth Q network algorithm to determine the optimal action space vector based on the current state space vector and perform full-level balancing control of the energy storage battery system.