Battery pooling scheduling method and device based on reinforcement learning

By employing a battery pooling scheduling method based on reinforcement learning, and combining battery cabinet operation data and perturbation response data, the problems of boundary risk and aging accumulation in multi-battery cabinet systems are solved. This enables comprehensive scheduling and recovery response of battery cabinets, thereby improving the operational capability and economy of energy storage systems.

CN122456701APending Publication Date: 2026-07-24浙江达航数据技术有限公司 +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
浙江达航数据技术有限公司
Filing Date
2026-04-22
Publication Date
2026-07-24

Smart Images

  • Figure CN122456701A_ABST
    Figure CN122456701A_ABST
Patent Text Reader

Abstract

The present application relates to the field of battery dispatching control, and particularly relates to a battery pooling dispatching method and device based on reinforcement learning, which comprises the following steps: obtaining operation data and perturbation response data of each battery cabinet, and determining cabinet-level capability state parameters; determining a multi-boundary risk index, an accumulated aging cost and a first boundary triggering time based on the cabinet-level capability state parameters; determining power allocation parameters and recovery allocation parameters in combination with a target power task, and generating a pooling dispatching instruction through reinforcement learning dispatching and operation constraint verification; when the first boundary triggering time is earlier than a target boundary time, performing energy reallocation, and updating related parameters according to a recovery response. The present application can improve dispatching continuity, boundary safety and life utilization balance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of battery scheduling and control, and more specifically to a battery pooling scheduling method and apparatus based on reinforcement learning. Background Technology

[0002] Solid-state battery energy storage systems possess high energy density and high safety potential. However, during long-term operation, the release capacity, temperature rise status, individual cell voltage margin, and interface stability of different battery cabinets will gradually diverge. Existing scheduling methods mostly allocate power based on voltage, state of charge, or fixed rules, making it difficult to take into account boundary risks, aging accumulation, and recovery needs. This can easily lead to some battery cabinets reaching their operating limits prematurely, insufficient release of available capacity, and uneven utilization of lifespan, thereby affecting the continuous operation capability, regulation accuracy, and overall life-cycle economics of the energy storage system. Summary of the Invention

[0003] This invention provides a battery pooling scheduling method and apparatus based on reinforcement learning, which is used to at least solve the problem of how to perform continuous pooling scheduling while taking into account boundary risks, aging accumulation and recovery requirements under multi-battery cabinet conditions.

[0004] In a first aspect, the present invention provides a battery pooling scheduling method based on reinforcement learning, the method comprising: Obtain the operation data and perturbation response data of each battery cabinet to determine the cabinet-level capability status parameters of each battery cabinet; Based on the capability status parameters of each cabinet, determine the multi-boundary risk indicators, cumulative aging costs, and the first boundary trigger time for each battery cabinet; Based on various multi-boundary risk indicators, cumulative aging costs, the first boundary trigger time, and the target power task, the power allocation parameters and recovery allocation parameters of each battery cabinet are determined. Based on these parameters, a reinforcement learning scheduler is used to generate the scheduling role and power share of each battery cabinet. The generated results are then subjected to operational constraint verification to obtain pooled scheduling instructions. In response to the first boundary trigger time after the execution of the pooled scheduling instruction being earlier than the target boundary time corresponding to the target power task, the energy redistribution path is determined based on the optimal transport method and energy redistribution is performed. Based on the recovery response after energy redistribution, the capacity status parameters of each cabinet level, multi-boundary risk indicators and cumulative aging costs are updated, and the pooled scheduling instruction for the next scheduling cycle is output.

[0005] In one possible implementation, the operation data and perturbation response data of each battery cabinet are acquired to determine the cabinet-level capability status parameters of each battery cabinet. This includes: acquiring the operation data of each battery cabinet, which includes the total voltage of the battery cabinet, the total current of the battery cabinet, the highest voltage of each cell, the lowest voltage of each cell, temperature data, and insulation status; acquiring the perturbation response data of each battery cabinet, which includes instantaneous voltage drop data, recovery slope data, and temperature response data; and determining the cabinet-level capability status parameters of each battery cabinet based on the operation data and perturbation response data.

[0006] In one possible implementation, cabinet-level capability state parameters include releasable energy margin, absorbable energy margin, thermal safety margin, individual cell voltage constraint margin, and interface stability margin.

[0007] In one possible implementation, based on the capability state parameters of each cabinet level, the multi-boundary risk indicators and the first boundary trigger time of each battery cabinet are determined, including: determining the charging boundary risk indicator, discharging boundary risk indicator, temperature boundary risk indicator, single-cell voltage boundary risk indicator, and interface stability boundary risk indicator based on the capability state parameters of each cabinet level; determining the boundary margin sequence of each battery cabinet based on the charging boundary risk indicator, discharging boundary risk indicator, temperature boundary risk indicator, single-cell voltage boundary risk indicator, and interface stability boundary risk indicator; and determining the first moment in the boundary margin sequence that is lower than a preset boundary threshold as the first boundary trigger time of each battery cabinet.

[0008] In one possible implementation, determining the cumulative aging cost of each battery cabinet includes: determining the aging increment of each battery cabinet based on the power load result and historical operating status of each battery cabinet in the current scheduling cycle; determining the aging cost reduction of each battery cabinet based on the boundary margin recovery amount of each battery cabinet in the recovery phase; and updating the cumulative aging cost of each battery cabinet based on each aging increment and each aging cost reduction.

[0009] In one possible implementation, based on various multi-boundary risk indicators, cumulative aging costs, the first boundary trigger time, and the target power task, the power allocation parameters and recovery allocation parameters of each battery cabinet are determined, including: generating a set of candidate roles for each battery cabinet based on various multi-boundary risk indicators, cumulative aging costs, and the first boundary trigger time. The set of candidate roles includes a set of candidate discharge roles, a set of candidate charging roles, a set of candidate buffer roles, a set of candidate recovery roles, a set of candidate observation roles, and a set of candidate isolation roles; and determining the power allocation parameters and recovery allocation parameters of each battery cabinet according to each set of candidate roles and the target power task.

[0010] In one possible implementation, based on power allocation parameters and recovery allocation parameters, a reinforcement learning scheduler is used to generate the scheduling role and power share of each battery cabinet, and a pooled scheduling instruction is obtained under the condition of satisfying operational constraints. This includes: inputting each power allocation parameter, recovery allocation parameter, role candidate set, and target power task into the reinforcement learning scheduler to generate candidate scheduling schemes; evaluating each candidate scheduling scheme according to each initial boundary trigger time and each cumulative aging cost to determine the target scheduling scheme; performing operational constraint verification on the target scheduling scheme; and correcting the target scheduling scheme if the operational constraint verification fails to pass, thereby obtaining the pooled scheduling instruction. The operational constraints include individual cell voltage constraints, battery cabinet current constraints, temperature constraints, insulation constraints, and contactor constraints.

[0011] In one possible implementation, determining the energy redistribution path and performing energy redistribution based on the optimal transport method includes: determining the source battery cabinet and the target battery cabinet based on the deviation between the initial boundary trigger time and the target boundary time corresponding to the target power task; determining the transport cost based on energy transfer loss, temperature rise, and the increase in the initial boundary trigger time after energy redistribution; determining the energy redistribution path based on the transport cost, and performing the energy redistribution corresponding to the energy redistribution path when the load fluctuation is lower than a preset threshold.

[0012] In one possible implementation, the capacity status parameters, multi-boundary risk indicators, and cumulative aging costs of each cabinet level are updated based on the recovery response after energy redistribution. This includes: collecting the recovery response after energy redistribution and determining the boundary margin recovery amount; updating the capacity status parameters, multi-boundary risk indicators, and cumulative aging costs of each cabinet level based on the boundary margin recovery amount; comparing the recovery response after energy redistribution with the predicted recovery response before energy redistribution; and correcting the power allocation parameters and recovery allocation parameters if the difference obtained from the comparison exceeds a preset threshold.

[0013] Secondly, the present invention provides a battery pooling scheduling apparatus based on reinforcement learning for implementing a battery pooling scheduling method based on reinforcement learning. The apparatus includes: The status determination module is used to acquire the operating data and perturbation response data of each battery cabinet and determine the cabinet-level capability status parameters of each battery cabinet. The risk determination module is used to determine the multi-boundary risk indicators, cumulative aging costs, and first boundary trigger time for each battery cabinet based on the capacity status parameters of each cabinet level. The scheduling generation module is used to determine the power allocation parameters and recovery allocation parameters of each battery cabinet based on various multi-boundary risk indicators, cumulative aging costs, the first boundary trigger time, and the target power task. Based on the power allocation parameters and recovery allocation parameters, it uses a reinforcement learning scheduler to generate the scheduling role and power share of each battery cabinet, performs operational constraint verification on the generated results, and obtains pooled scheduling instructions. The redistribution module is used to respond when the first boundary trigger time after the execution of the pooled scheduling instruction is earlier than the target boundary time corresponding to the target power task. It determines the energy redistribution path based on the optimal transport method and executes the energy redistribution. Based on the recovery response after energy redistribution, it updates the capacity status parameters of each cabinet level, multi-boundary risk indicators and cumulative aging costs, and outputs the pooled scheduling instruction for the next scheduling cycle.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: By employing a joint characterization technique combining operational data and perturbation response data, the true capability status of the battery cabinet was identified, avoiding distortion caused by allocating power solely based on static voltage. Through joint modeling techniques involving multiple boundary risks, cumulative aging costs, and the initial boundary trigger moment, the scheduling basis was transformed from a single state variable to comprehensive boundary constraints. By using a collaborative technique combining reinforcement learning scheduling and operational constraint verification, the unification of power allocation and recovery allocation was achieved. Through energy redistribution and recovery response write-back techniques, the repair of bottleneck battery cabinets and rolling correction for the next cycle were realized. Attached Figure Description

[0015] Figure 1 This is a schematic flowchart of the method of the present invention; Figure 2 This is a multi-boundary risk distribution diagram in a specific embodiment of the present invention; Figure 3 This is a target power and power sharing diagram in a specific embodiment of the present invention; Figure 4 This is a comparison diagram of the first boundary triggering time in a specific embodiment of the present invention; Figure 5 This is a recovery response curve after energy redistribution in a specific embodiment of the present invention; Figure 6 This is a block diagram of the modular components of the device of the present invention. Detailed Implementation

[0016] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0019] Reinforcement learning is a sequential decision-making technique based on interactive feedback. Its core lies in the continuous iteration between state perception, action selection, and benefit evaluation to gradually form a control strategy suitable for dynamic environments. Compared to control methods that rely solely on fixed rules or static mapping relationships, reinforcement learning is better suited for handling scheduling problems with multi-objective constraints, temporal correlations, and long-term benefit trade-offs. It can comprehensively consider the current state, subsequent impacts, and overall control costs during continuous decision-making. For scheduling tasks involving the coordinated operation of multiple battery cabinets, the capacity states, boundary constraints, and recovery requirements of different battery cabinets vary significantly at different times. Allocation based solely on a single indicator is insufficient to balance overall consistency and continuity. Therefore, it is necessary to introduce a temporally-oriented decision-making scheduling mechanism, forming a pooled scheduling method that can be dynamically updated during operation, based on integrated state identification, boundary judgment, and task allocation. Based on this, a battery pooled scheduling method based on reinforcement learning is proposed.

[0020] like Figure 1 As shown, a battery pooling scheduling method based on reinforcement learning is proposed, which includes: Obtain the operation data and perturbation response data of each battery cabinet to determine the cabinet-level capability status parameters of each battery cabinet; After acquiring the operating data and perturbation response data of each battery cabinet, the data from different acquisition channels are first time-aligned and their validity is checked. Then, the voltage, current, temperature, and insulation status within the same acquisition cycle are combined into a battery cabinet operating status segment. When the battery cabinet is in a state of small power fluctuation and has not reached the safety boundary, a test perturbation with limited amplitude is applied to the corresponding battery cabinet, and the voltage change, recovery process, and temperature change before and after the perturbation are recorded to form perturbation response data. Subsequently, the operating status segment and the perturbation response data are correlated and analyzed to extract state features characterizing energy margin, temperature tolerance, individual cell voltage approximation degree, and interface contact stability, and obtain cabinet-level capability state parameters for subsequent risk assessment and scheduling allocation.

[0021] The process involves acquiring operational data and perturbation response data for each battery cabinet to determine its cabinet-level capability status parameters. This includes: acquiring operational data for each battery cabinet, including total voltage, total current, maximum voltage of individual cells, minimum voltage of individual cells, temperature data, and insulation status; acquiring perturbation response data for each battery cabinet, including instantaneous voltage drop data, recovery slope data, and temperature response data; and determining the cabinet-level capability status parameters for each battery cabinet based on the operational data and perturbation response data.

[0022] In one embodiment, operational data acquisition employs a combination of periodic sampling and event sampling. Periodic sampling continuously records the total voltage and current of the battery cabinet, the highest and lowest voltage of individual cells, temperature data, and insulation status. Event sampling increases the sampling density during charging and discharging transitions, power transitions, rapid temperature changes, and before and after protection actions. This approach is chosen because cabinet-level capability status parameters require both long-term stable operating trajectories and the ability to reflect instantaneous characteristics during rapid state changes.

[0023] The total voltage and current of the battery cabinet characterize the current energy exchange level. The highest and lowest voltages of individual cells reflect the dispersion of individual cells and whether they are close to voltage boundaries. Temperature data reflects the heat load distribution, and insulation status determines whether the current battery cabinet is suitable for subsequent scheduling calculations. Temperature data is preferably generated from multiple temperature monitoring points located at the top, middle, bottom, and weak heat dissipation areas of the battery cabinet. This avoids misjudgments based solely on single-point temperatures. For acquiring perturbation response data, no additional high-power excitation is required. Instead, small-scale test perturbations are performed during periods when load fluctuations are below a preset threshold, without affecting normal operation. The preset threshold can be set based on the stable operating range of the battery cabinet's rated power; for example, a small portion of the rated power can be selected as the allowable fluctuation range to ensure that the test perturbation does not overlap with external scheduling commands, creating a significant impact. Test perturbations can employ short-term power increases, short-term power decreases, or short-term zero-change followed by recovery observation.

[0024] Regardless of the method used, the duration is controlled within a time range that conventional battery management systems in this field can stably identify. During data acquisition, stable data before the disturbance begins is used as the baseline segment, data during the disturbance application period is used as the response segment, and data after the disturbance is removed is used as the recovery segment. Instantaneous voltage drop data is taken from the voltage change difference between the baseline and response segments, recovery slope data is taken from the rate of change of voltage recovery in the recovery segment, and temperature response data is taken from the temperature change trend and temperature difference diffusion process before and after the disturbance. To ensure data usability, after generating operational data and disturbance response data, further outlier removal, missing value completion, and timestamp correction can be performed. Outlier removal is preferably based on a combination of the sampling range and the change amplitude of adjacent time points, missing value completion is preferably achieved using interpolation of adjacent stable segments, and timestamp correction is preferably based on the cabinet-level controller clock. Through the above processing, a data foundation with a clear source, consistent time, and directly applicable to the state parameter solution process can be obtained.

[0025] Cabinet-level capability status parameters include release energy margin, absorbable energy margin, thermal safety margin, individual cell voltage constraint margin, and interface stability margin.

[0026] In one embodiment, the cabinet-level capability state parameters include releasable energy margin, absorbable energy margin, thermal safety margin, individual cell voltage constraint margin, and interface stability margin. The releasable energy margin represents the capability boundary of the battery cabinet to continue performing discharge tasks in the current state; the absorbable energy margin represents the capability boundary of the battery cabinet to continue performing charging tasks in the current state; the thermal safety margin represents the remaining space of the battery cabinet from the temperature limit boundary under a given task intensity; the individual cell voltage constraint margin represents the remaining space of the highest and lowest individual cell voltages from the allowable boundaries; and the interface stability margin represents the stability of the internal contact state of the solid-state battery under the current stress level.

[0027] The parameters mentioned above are not directly given by a single sampled value, but are determined based on operational data and perturbation response data. For the release and absorption energy margins, the total battery cabinet voltage, total battery cabinet current, and their changing trends over continuous sampling periods can be considered, and the overall energy boundary can be corrected by referring to the highest and lowest individual cell voltages. This is because, when a solid-state battery cabinet approaches its boundary, individual cell differences will appear before the total quantity. For the thermal safety margin, it can be determined based on the current temperature distribution, maximum temperature difference, and temperature response data. If the temperature drops slowly after a test perturbation, or if the temperature difference between adjacent detection points widens significantly, it indicates that the battery cabinet's thermal tolerance is weak under the current condition, and the power load ratio in subsequent scheduling needs to be appropriately reduced. For the individual cell voltage constraint margin, it can be obtained by the difference between the highest individual cell voltage and the upper charging limit, and the difference between the lowest individual cell voltage and the lower discharging limit, respectively, and the tighter portion is taken as the characterization result of the individual cell voltage constraint margin.

[0028] For interface stability margin, it can be judged by combining instantaneous voltage drop data and recovery slope data. When the voltage recovery process is smooth and the recovery speed remains within the historical normal range after the test disturbance is removed, the current interface contact state can be judged to be relatively stable; when the instantaneous voltage drop increases significantly, the recovery process trails, or the recovery speed continues to decrease under the same disturbance conditions, the current interface stability can be judged to be reduced. The principle for setting the interface stability margin can be based on the benchmark response range formed during the historical stable operation phase, or it can be based on the performance boundary given by the manufacturer. This ensures that the parameter reflects the operating characteristics of the battery cabinet itself, while not deviating from the actual executable conditions. After the above parameters are determined, each parameter can be encapsulated into the corresponding state description result of the battery cabinet and bound to the acquisition time for subsequent multi-boundary risk index calculation, cumulative aging cost update, and scheduling role allocation.

[0029] Based on the capability status parameters of each cabinet, determine the multi-boundary risk indicators, cumulative aging costs, and the first boundary trigger time for each battery cabinet; After obtaining the cabinet-level capability status parameters of each battery cabinet, risk assessment results are first established according to five types of operational boundaries: charging, discharging, temperature, individual cell voltage, and interface stability. These five risk assessment results are then uniformly mapped to a boundary margin sequence within the same scheduling cycle. Subsequently, the boundary margin sequence is retrieved hourly along the prediction time axis, and the moment when the boundary margin first falls below a preset boundary threshold is determined as the first boundary trigger moment for the corresponding battery cabinet. Simultaneously, combining the power load of each battery cabinet within the current scheduling cycle, historical operating status, and boundary margin recovery during the recovery phase, the aging increment and aging cost reduction of each battery cabinet in the current cycle are calculated, and the cumulative aging cost is updated accordingly. After this processing, three types of results reflecting the current risk level, future boundary approach, and long-term consumption level can be obtained simultaneously, serving as direct inputs for subsequent power allocation and recovery allocation.

[0030] Based on the capability status parameters of each cabinet level, the multi-boundary risk indicators and the first boundary trigger time of each battery cabinet are determined, including: based on the capability status parameters of each cabinet level, determining the charging boundary risk indicator, discharging boundary risk indicator, temperature boundary risk indicator, single-cell voltage boundary risk indicator, and interface stability boundary risk indicator respectively; based on the charging boundary risk indicator, discharging boundary risk indicator, temperature boundary risk indicator, single-cell voltage boundary risk indicator, and interface stability boundary risk indicator, determining the boundary margin sequence of each battery cabinet; and determining the first moment in the boundary margin sequence that is lower than the preset boundary threshold as the first boundary trigger time of each battery cabinet.

[0031] In one embodiment, the determination of multiple boundary risk indicators and the initial boundary trigger time adopts a "single boundary judgment, unified conversion, and sequential retrieval" processing method. First, charging boundary risk indicators, discharging boundary risk indicators, temperature boundary risk indicators, single-cell voltage boundary risk indicators, and interface stability boundary risk indicators are established for each battery cabinet. The charging boundary risk indicator is mainly determined based on the absorbable energy margin, the remaining space between the single-cell maximum voltage and the charging upper limit, and the recent temperature change trend; when the absorbable energy margin continues to decrease, the single-cell maximum voltage approaches the charging upper limit, or the temperature rise rate during charging significantly accelerates, the charging boundary risk indicator increases accordingly.

[0032] The discharge boundary risk index is mainly determined based on the releaseable energy margin, the remaining space between the minimum voltage of a single cell and the lower discharge limit, and the voltage decline trend during continuous discharge. When the releaseable energy margin is insufficient and the minimum voltage of a single cell continues to approach the lower discharge limit, the discharge boundary risk index increases accordingly. The temperature boundary risk index is determined based on the thermal safety margin, the maximum temperature difference, and its expansion trend. When the highest temperature approaches the upper limit of the allowable temperature, or the temperature difference between different measuring points continues to widen, it indicates uneven heat distribution inside the battery cabinet, and the temperature boundary risk index increases accordingly. The single cell voltage boundary risk index is determined based on the highest voltage of a single cell, the lowest voltage of a single cell, and the fluctuation range of both within the current cycle. When both the highest and lowest single cell voltages show a trend of approaching the boundary, the side closer to the boundary is preferred as the risk judgment basis. The interface stability boundary risk index is determined based on the interface stability margin, the instantaneous voltage drop change, and the smoothness of the recovery process. If the instantaneous voltage drop increases and the recovery process is prolonged under the same disturbance conditions, the interface stability is judged to have decreased. After completing the assessment of the five types of risks, the various risk indicators are converted into a boundary margin sequence in a unified direction, so that the smaller the boundary margin value, the closer it is to the corresponding boundary.

[0033] The preset boundary threshold is preferably determined based on stable operating samples under rated conditions, or it can be set by shrinking the safety control boundary provided by the manufacturer by a certain proportion to ensure sufficient safety margin for subsequent scheduling. Then, the values ​​of each moment in the boundary margin sequence are compared sequentially along the prediction time axis, and the first moment that falls below the preset boundary threshold is determined as the first boundary trigger moment. If no moment falls below the preset boundary threshold within the current prediction window, the estimated moment after the end of the prediction window is recorded as the extrapolated result of the first boundary trigger moment, indicating that the current battery cabinet still has a schedulable margin within this window. After this processing, constraints from different physical boundaries are transformed into a unified time quantity, facilitating direct comparison of the boundary arrival order of different battery cabinets under the same target power task.

[0034] The cumulative aging cost of each battery cabinet is determined, including: determining the aging increment of each battery cabinet based on the power load results and historical operating status of each battery cabinet in the current scheduling cycle; determining the aging cost reduction of each battery cabinet based on the boundary margin recovery amount of each battery cabinet in the recovery phase; and updating the cumulative aging cost of each battery cabinet based on each aging increment and each aging cost reduction.

[0035] In one embodiment, the cumulative aging cost is determined using a combination of "cycle increment calculation and recovery correction". First, the corresponding aging increment is determined based on the power load of each battery cabinet within the current scheduling cycle and its historical operating status. The power load results include at least the charging power, discharging power, duration, and frequency of power changes within the current cycle; the historical operating status includes at least the average power level, temperature burden level, individual cell voltage approximation, and accumulated aging cost from previous scheduling cycles. For battery cabinets with higher charging or discharging power, longer durations, more frequent power changes, and heavier temperature burdens, the aging increment is set larger; for battery cabinets with lighter power loads in the current cycle and lower historical burdens, the aging increment is set smaller. This is because the long-term consumption of a battery cabinet is not only related to the magnitude of a single power input, but also to the duration, the degree of power change, and the accumulated operating pressure from previous cycles.

[0036] Subsequently, the reduction in aging costs is determined based on the increase in boundary margin during the recovery phase. The recovery phase can be a low-load phase, a static phase, or a recovery phase specifically allocated in subsequent scheduling after the end of the current scheduling cycle. The increase in boundary margin reflects the degree to which various operating boundaries are reopened after the battery cabinet is removed from a high load. If, after the recovery phase, the energy release margin, absorbable energy margin, thermal safety margin, or interface stability margin show a significant increase, it indicates that the consumption incurred by the battery cabinet in this cycle has been partially released, and the reduction in aging costs can be appropriately increased. If the increase in boundary margin is not significant after the recovery phase, it indicates that the battery cabinet is still under high pressure, and the reduction in aging costs should be kept small. The principle for setting the reduction in aging costs is preferably to maintain a monotonic correspondence with the increase in boundary margin, that is, the more significant the increase in boundary margin, the greater the reduction in aging costs, but it should not exceed the aging increment of this cycle to avoid abnormal reverse fluctuations in cumulative aging costs.

[0037] Finally, the aging increment of the current cycle and the reduction in aging cost are combined and applied to the cumulative aging cost of the previous cycle to obtain the updated cumulative aging cost. If the updated result is lower than a preset lower limit, the cumulative aging cost is corrected to the preset lower limit to avoid overly optimistic judgments about the state of the battery cabinet in subsequent scheduling due to short-term recovery. Through the above processing, the cumulative aging cost retains the long-term consumption characteristics caused by historical operation and absorbs the state mitigation effect brought about by the recovery phase, which can more realistically reflect the cost level of each battery cabinet continuing to undertake power tasks in subsequent scheduling.

[0038] Based on various multi-boundary risk indicators, cumulative aging costs, the first boundary trigger time, and the target power task, the power allocation parameters and recovery allocation parameters of each battery cabinet are determined. Based on these parameters, a reinforcement learning scheduler is used to generate the scheduling role and power share of each battery cabinet. The generated results are then subjected to operational constraint verification to obtain pooled scheduling instructions. After obtaining multi-boundary risk indicators, cumulative aging costs, and the initial boundary trigger time, the target power task is first decomposed into charging demand, discharging demand, power fluctuation tolerance demand, and recovery demand. Then, based on the current boundary position and historical consumption level of each battery cabinet, a set of role candidates is generated for each battery cabinet. Subsequently, according to the matching relationship between the role candidate set and the target power task, the power allocation parameters and recovery allocation parameters are determined respectively. Then, the power allocation parameters and recovery allocation parameters are input into the reinforcement learning scheduler to obtain the scheduling role and power share of each battery cabinet. The generated results are then subjected to operational constraint verification based on individual cell voltage, battery cabinet current, temperature, insulation status, and contactor status to form pooled scheduling instructions that can be directly issued.

[0039] Based on various multi-boundary risk indicators, cumulative aging costs, the first boundary trigger time, and the target power task, the power allocation parameters and recovery allocation parameters of each battery cabinet are determined, including: generating a set of candidate roles for each battery cabinet based on various multi-boundary risk indicators, cumulative aging costs, and the first boundary trigger time. The set of candidate roles includes a set of candidate discharge roles, a set of candidate charging roles, a set of candidate buffer roles, a set of candidate recovery roles, a set of candidate observation roles, and a set of candidate isolation roles; and determining the power allocation parameters and recovery allocation parameters of each battery cabinet based on each set of candidate roles and the target power task.

[0040] In one embodiment, the determination of power allocation parameters and recovery allocation parameters adopts a three-stage process: "candidate role screening, task component matching, and quota calculation." First, a candidate set of roles is established for each battery cabinet based on multiple boundary risk indicators, cumulative aging costs, and the initial boundary trigger time. The candidate set of roles includes at least a discharge role candidate set, a charging role candidate set, a buffer role candidate set, a recovery role candidate set, an observation role candidate set, and an isolation role candidate set. The discharge role candidate set is used to accommodate battery cabinets with continuous output capabilities and relatively late boundary trigger times; the charging role candidate set is used to accommodate battery cabinets with continuous absorption capabilities and sufficient margin between the highest voltage of individual cells and the charging boundary; the buffer role candidate set is used to accommodate battery cabinets suitable for handling short-term power fluctuations but not suitable for long-term high loads; the recovery role candidate set is used to accommodate battery cabinets with high cumulative aging costs and strong boundary recovery requirements; the observation role candidate set is used to accommodate battery cabinets with a certain boundary approach trend but have not yet reached the isolation condition; and the isolation role candidate set is used to accommodate battery cabinets that are currently not suitable for participating in scheduling and allocation.

[0041] The roles mentioned above are not fixed assignments, but are regenerated at the beginning of each scheduling cycle. Specifically, during generation, each battery cabinet can be initially sorted according to its first boundary trigger time, and then adjusted based on accumulated aging costs. If a battery cabinet has a later first boundary trigger time and lower accumulated aging costs, it is prioritized for inclusion in the discharge or charging role candidate set; if its first boundary trigger time is in the middle but its temperature boundary risk index or interface stability boundary risk index is high, it is prioritized for inclusion in the buffer role candidate set; if its accumulated aging costs are consistently higher than the average level of the same group, or if the boundary margin recovery during the recovery phase is insufficient, it is prioritized for inclusion in the recovery role candidate set; if the insulation condition is abnormal, the contactor operation is unstable, or the individual cell voltage boundary risk index remains high, it is included in the observation or isolation role candidate set. After the role candidate sets are determined, the target power task is further divided into a base power component, a fluctuation regulation component, and a recovery retention component.

[0042] The base power component is preferentially allocated to battery cabinets in the discharge or charging candidate sets, the fluctuation regulation component is preferentially allocated to battery cabinets in the buffer candidate set, and the recovery retention component is preferentially allocated to battery cabinets in the recovery candidate set. Power allocation parameters can include allocation direction, target power range, and duration. Recovery allocation parameters can include recovery duration, allowable power limit, and conditions for re-participation in scheduling. The principle for setting the recovery duration can be determined by combining the cumulative aging cost level and the recovery effect of the previous cycle; the higher the cumulative aging cost and the slower the boundary recovery in the previous cycle, the longer the recovery duration. The allowable power limit can be set based on the greater of the temperature boundary risk index and the individual cell voltage boundary risk index, thereby avoiding the introduction of excessive burden during the recovery phase. Through the above method, multiple boundary risks, long-term consumption, and current task requirements can be uniformly mapped to executable power allocation parameters and recovery allocation parameters.

[0043] Based on the power allocation parameters and recovery allocation parameters, a reinforcement learning scheduler is used to generate the scheduling role and power share of each battery cabinet. Under the condition of satisfying the operational constraints, a pooled scheduling instruction is obtained. This includes: inputting the power allocation parameters, recovery allocation parameters, role candidate set, and target power task into the reinforcement learning scheduler to generate candidate scheduling schemes; evaluating each candidate scheduling scheme according to each initial boundary trigger time and each cumulative aging cost to determine the target scheduling scheme; performing operational constraint verification on the target scheduling scheme; and correcting the target scheduling scheme if the operational constraint verification fails to pass, thereby obtaining the pooled scheduling instruction. The operational constraints include individual cell voltage constraints, battery cabinet current constraints, temperature constraints, insulation constraints, and contactor constraints.

[0044] In one embodiment, the generation of scheduling roles and power shares follows a sequence of "candidate scheme generation, scheme evaluation, constraint verification, and result correction." After receiving power allocation parameters, recovery allocation parameters, a set of candidate roles, and the target power task, the reinforcement learning scheduler does not directly output a unique allocation result. Instead, it first generates multiple sets of candidate scheduling schemes. Each set of candidate scheduling schemes includes at least the scheduling role of each battery cabinet, its corresponding power share, and a recovery continuity arrangement. In addition to the power allocation parameters and recovery allocation parameters generated in the current cycle, the reinforcement learning scheduler's input state can also include the actual execution result from the previous cycle to maintain scheduling continuity.

[0045] After candidate scheduling schemes are generated, they need to be evaluated in conjunction with the initial boundary trigger time and cumulative aging costs. During the evaluation, schemes that significantly advance the boundary trigger time are prioritized for elimination, followed by schemes that further concentrate cumulative aging costs on a few battery cabinets. From the remaining schemes, the one that meets the target power requirement and has sufficient recovery arrangements is selected as the target scheduling scheme. To ensure that the evaluation results can be directly applied to field execution, the evaluation process does not rely on complex and difficult-to-implement high-order solutions, but instead adopts a verification method directly corresponding to the operational boundaries. For example, each cabinet can be checked to see if the power allocation in the target scheduling scheme will cause the highest voltage of a single cell to exceed the charging boundary, whether it will cause the lowest voltage of a single cell to fall below the discharge boundary, whether it will cause the battery cabinet current to exceed the rated range, whether it will cause the temperature to reach the preset upper limit, and whether it will continue to allocate power under conditions of abnormal insulation or contactor status. After completing the above checks, the target scheduling scheme is verified for operational constraints. Operational constraints include single cell voltage constraints, battery cabinet current constraints, temperature constraints, insulation constraints, and contactor constraints.

[0046] If the operational constraint verification passes, the target scheduling scheme is directly transformed into a pooled scheduling instruction. If the operational constraint verification fails, the target scheduling scheme is modified. During modification, the power share of the battery cabinet that triggered the constraint is reduced first, and the reduced portion is redistributed to battery cabinets with later boundary trigger times and lower cumulative aging costs in the candidate set of similar roles. If there are no replacement battery cabinets in the candidate set of similar roles, the corresponding battery cabinet is adjusted to a buffer role or a recovery role, and the execution slope of the corresponding component in the target power task is reduced until the operational constraints are met. For battery cabinets with abnormal insulation or contactor states, the current cycle power allocation is directly canceled, and only observation or isolation arrangements are retained. After modification, the operational constraint verification is performed again until a pooled scheduling instruction that meets the current operational boundary and can cover the requirements of the target power task is formed. After this processing, the output of the reinforcement learning scheduler no longer remains at the abstract allocation level, but can be transformed into scheduling roles and power shares that can be directly executed by each battery cabinet.

[0047] In response to the first boundary trigger time after the execution of the pooled scheduling instruction being earlier than the target boundary time corresponding to the target power task, the energy redistribution path is determined based on the optimal transport method and energy redistribution is performed. Based on the recovery response after energy redistribution, the capacity status parameters of each cabinet level, multi-boundary risk indicators and cumulative aging costs are updated, and the pooled scheduling instruction for the next scheduling cycle is output.

[0048] In this embodiment, after the pooled scheduling instruction completes the current scheduling cycle, the initial boundary trigger time of each battery cabinet is compared with the target boundary time corresponding to the target power task. When the initial boundary trigger time is earlier than the target boundary time, it indicates that relying solely on the current power allocation and recovery allocation is insufficient to support subsequent tasks. At this time, the battery cabinets participating in the scheduling are re-screened at the source and target ends to determine the energy redistribution path, and energy redistribution is implemented when the execution conditions are met. After the energy redistribution is completed, the voltage, temperature, and boundary margin changes during the recovery phase are collected, and the cabinet-level capability status parameters, multi-boundary risk indicators, and cumulative aging costs are updated. Based on this, the pooled scheduling instruction for the next scheduling cycle is generated, ensuring that subsequent scheduling maintains task continuity while preventing individual battery cabinets from prematurely reaching the operating boundary.

[0049] Determining and executing energy redistribution based on the optimal transport method includes: determining the source battery cabinet and the target battery cabinet based on the deviation between the initial boundary trigger time and the target boundary time corresponding to the target power task; determining the transport cost based on energy transfer loss, temperature rise, and the increase in the initial boundary trigger time after energy redistribution; determining the energy redistribution path based on the transport cost, and executing the energy redistribution corresponding to the energy redistribution path when the load fluctuation is lower than a preset threshold.

[0050] In one embodiment, the determination of the energy redistribution path follows the sequence of "deviation identification, path screening, cost comparison, and window execution." First, battery cabinets are grouped based on the deviation between their initial boundary trigger time and the target boundary time corresponding to the target power task. Battery cabinets whose initial boundary trigger time is significantly earlier than the target boundary time are assigned to the target battery cabinet set; battery cabinets whose initial boundary trigger time is later than the target boundary time but whose available energy margin is still within the allowable range are assigned to the source battery cabinet set. This division is based on the fact that target battery cabinets require additional energy support to delay the boundary arrival time, while source battery cabinets need to relinquish some available energy without affecting their subsequent tasks. After screening source and target battery cabinets, candidate redistribution paths are constructed one by one. Each candidate redistribution path corresponds to at least one source battery cabinet and one target battery cabinet, and the expected energy transfer loss, expected temperature rise, and expected boundary improvement under that path are recorded.

[0051] To facilitate a unified comparison among multiple candidate paths, a transportation cost can be calculated for each candidate path. The transportation cost can be expressed as:

[0052] in, Power source battery cabinet To the target battery cabinet The transport costs during energy redistribution; This represents the expected energy transfer loss for this path; This represents the expected temperature rise along this path; This represents the expected boost at the moment of the first boundary trigger after the path is executed; This is the weighting coefficient for energy transfer losses; This is the temperature rise weighting coefficient; The weighting coefficients for boundary improvement are set. The weighting coefficients for energy transfer loss, temperature rise, and boundary improvement can be set based on historical operating data during the commissioning phase.

[0053] When the system prioritizes conversion efficiency, the weighting coefficient for energy transfer loss is increased; when the system prioritizes thermal safety, the weighting coefficient for temperature rise is increased; and when the system prioritizes the arrival time of delayed boundaries, the weighting coefficient for boundary improvement is increased. After calculating the transport cost, the path with the lowest transport cost and where neither the source nor the target battery cabinet touches the current safety boundary is selected as the energy redistribution path from the candidate redistribution paths. To avoid the energy redistribution action overlapping with drastic fluctuations in external load, it is also necessary to determine whether the current load fluctuation is below a preset threshold. The preset threshold can be set according to a certain proportion of the rated power, with the principle being that the additional power during energy redistribution will not cause the station-level power to deviate significantly from the target power task. When the load fluctuation is below the preset threshold, the energy redistribution of the corresponding energy redistribution path is executed; when the load fluctuation is above the preset threshold, the current pooling scheduling instruction is maintained, and the energy redistribution task is postponed to the next executable window. This ensures that the timing of energy redistribution is consistent with the station-level operating status, avoiding the introduction of additional unstable factors when the load changes drastically.

[0054] The system updates the capacity status parameters, multi-boundary risk indicators, and cumulative aging costs of each cabinet level based on the recovery response after energy redistribution. This includes: collecting the recovery response after energy redistribution and determining the boundary margin recovery amount; updating the capacity status parameters, multi-boundary risk indicators, and cumulative aging costs of each cabinet level based on the boundary margin recovery amount; comparing the recovery response after energy redistribution with the predicted recovery response before energy redistribution; and correcting the power allocation parameters and recovery allocation parameters if the difference obtained from the comparison exceeds a preset threshold.

[0055] In one embodiment, after energy redistribution is completed, the state parameters before redistribution are not directly used for continued scheduling. Instead, the state model is backfilled and corrected through the recovery response. The recovery response refers to the voltage recovery process, temperature drop process, and boundary margin change process collected within a preset recovery observation window after energy redistribution. The preset recovery observation window can be set according to the battery cabinet voltage recovery rate and temperature change rate, generally covering the short-term stable phase after energy redistribution to ensure that the collected results can truly reflect the state changes brought about by the redistribution action.

[0056] First, operational data within the recovery observation window is collected to calculate the boundary margin recovery amount. The boundary margin recovery amount reflects the degree to which the target battery cabinet is further separated from the charging boundary, discharging boundary, temperature boundary, single-cell voltage boundary, and interface stability boundary after receiving energy compensation. A larger boundary margin recovery amount indicates a more significant effect of energy redistribution on delaying the boundary arrival time. Subsequently, the cabinet-level capability status parameters, multi-boundary risk indicators, and cumulative aging costs are updated based on the boundary margin recovery amount. For the target battery cabinet, the releaseable energy margin or absorbable energy margin is adjusted synchronously with the boundary margin recovery amount, and the multi-boundary risk indicators are recalculated according to the new boundary positions; for the source battery cabinet, the corresponding parameters are adjusted synchronously based on the boundary changes after energy transfer. The update of the cumulative aging costs should simultaneously consider the increased power burden during energy redistribution and the state mitigation brought about by the recovery phase. If the target battery cabinet shows a significant boundary margin recovery within the recovery observation window, the increase in cumulative aging costs is relatively small; if the source battery cabinet experiences a temperature increase or boundary approach due to energy output, the cumulative aging costs increase accordingly.

[0057] To avoid relying entirely on actual observations for state updates, the recovery response after energy redistribution is compared with the predicted recovery response before energy redistribution. The predicted recovery response can be provided by the state prediction results generated before the current scheduling cycle, reflecting the theoretically achievable recovery level. When the difference obtained from the comparison exceeds a preset threshold, it indicates that there is a deviation between the current power allocation parameters and the recovery allocation parameters in describing the actual state of the battery cabinet, requiring correction of the power allocation parameters and recovery allocation parameters for the next scheduling cycle. The preset threshold can be set according to the allowable error range of boundary margin changes within the recovery observation window. The setting principle is to identify significant deviations without frequently triggering corrections due to normal fluctuations. During correction, priority can be given to reducing the power load ratio of the battery cabinet corresponding to the deviation in the next scheduling cycle, or extending the recovery time of the corresponding battery cabinet. After completing the above updates and corrections, the pooled scheduling instructions for the next scheduling cycle are regenerated, so that subsequent scheduling is based on the latest state. After this processing, energy redistribution is not only used to temporarily fill task gaps, but also to reverse the state identification and parameter allocation results, forming a continuous closed loop.

[0058] In one specific embodiment, a six-cabinet parallel solid-state battery energy storage system was selected for verification. The system has a rated power of 800 kW and a rated capacity of 1.2 MWh, with each battery cabinet having a rated capacity of 200 kWh. The sampling period of the station-level controller was set to 1 second, the pooling scheduling period was set to 5 minutes, and the reinforcement learning scheduler updated the power share every 15 minutes. The test period was set from 2 PM to 4 PM, corresponding to one continuous discharge task. The target power tasks during this period were 620 kW, 700 kW, 760 kW, 780 kW, 720 kW, 680 kW, 640 kW, and 600 kW, respectively. The initial state of charge (SOC) of each battery cabinet was 72%, 69%, 61%, 58%, 74%, and 66%, respectively. Cabinets 3 and 4 had previously carried higher loads and had a higher risk of prematurely reaching the load limit.

[0059] Within 10 minutes before 2 PM, the station-level controller continuously collected operational data from six battery cabinets. Under conditions where load fluctuations were below 3% of rated power, a small-amplitude test disturbance was applied to each battery cabinet for 3 seconds. The operational data included total battery cabinet voltage, total battery cabinet current, highest and lowest individual cell voltages, temperature data, and insulation status. The disturbance response data included instantaneous voltage drop data, recovery slope data, and temperature response data. The processed cabinet-level capability status parameters showed that cabinets 1 and 5 had relatively high release energy margins of 178 kWh and 171 kWh, respectively; cabinets 3 and 4 had relatively low thermal safety margins of 0.61 and 0.58, respectively; and individual cell voltage constraint margins of only 66 mV and 61 mV, respectively, with interface stability margins of 0.59 and 0.55, respectively. Therefore, cabinets 3 and 4 are more suitable as protected objects than as primary power-bearing objects.

[0060] like Figure 2 As shown, Figure 2 This is a multi-boundary risk distribution map. The map was calculated based on operational and perturbation response data collected from the six battery cabinets 10 minutes before 2 PM. The horizontal axis represents the battery cabinet number, and the vertical axis represents the risk category. Higher values ​​indicate higher boundary risks. It can be seen that cabinets 3 and 4 have significantly higher discharge boundary risk, temperature boundary risk, single-cell voltage boundary risk, and interface stability boundary risk than the other battery cabinets. In particular, cabinet 4's interface stability boundary risk reaches 0.59, and its temperature boundary risk reaches 0.55, indicating that if the target power task continues to be distributed in the conventional manner, cabinet 4 is most likely to reach the operational boundary first. Cabinets 1 and 5 have relatively low overall risks across all categories and are suitable as the main energy output terminals.

[0061] After determining the multi-boundary risk indicators, the system further generates the initial boundary trigger time and cumulative aging cost. Based on the current state, the initial boundary trigger times for cabinets 1, 2, 3, 4, 5, and 6 are 150 minutes, 138 minutes, 118 minutes, 112 minutes, 146 minutes, and 129 minutes, respectively. Considering that the target boundary time corresponding to the target power task is set to 120 minutes, cabinets 3 and 4 obviously cannot stably complete the task. Regarding the cumulative aging cost, cabinets 3 and 4 have costs of 1.28 and 1.36, respectively, significantly higher than cabinet 1's 1.02 and cabinet 5's 1.05, further indicating that both should reduce power allocation and retain recovery space in the current cycle.

[0062] Based on this, the reinforcement learning scheduler assigns roles to the six battery cabinets. Cabinets 1, 2, and 5 are assigned to the primary discharge role, cabinet 6 to the buffer role, and cabinets 3 and 4 to the recovery priority role, while still retaining a small discharge share to ensure continuous fulfillment of the station-level target power. After scheduling, the power share allocation for each time period is as follows: Figure 3 As shown. Figure 3 The diagram shows the target power and power sharing. It is drawn based on the scheduling results every 15 minutes. The black line represents the station-level target power, and the different colored bars represent the power borne by each battery cabinet at the corresponding time. It can be seen that in the high-load range of 760 kW to 780 kW, cabinet 6 undertakes a significant buffering task, while the power borne by cabinets 3 and 4 is compressed to the range of 50 kW to 65 kW. Calculated cumulatively over 8 time periods, cabinet 1 borne 287.5 kWh, cabinet 2 borne 246.25 kWh, cabinet 3 borne 145 kWh, cabinet 4 borne 121.25 kWh, cabinet 5 borne 291.25 kWh, and cabinet 6 borne 283.75 kWh. Cabinets 3 and 4 combined only borne 266.25 kWh, accounting for 19.4% of the total discharge, indicating that this invention has removed high-risk battery cabinets from the primary power-bearing position.

[0063] Despite this, when the scheduling was re-evaluated at 3:15 PM, the initial boundary trigger times for cabinets 3 and 4 were still only 118 minutes and 112 minutes respectively, still earlier than the target boundary time of 120 minutes. At this point, the system triggered energy redistribution. To avoid additional disturbances, the redistribution window was selected from 3:20 PM to 3:30 PM, during which time the station-level load fluctuation was less than 15 kW. For candidate paths, transport costs were used for comparison:

[0064] in, This refers to the transport cost when energy is redistributed from the source battery cabinet to the target battery cabinet. For energy transfer losses, For temperature rise, This represents the boost amount at the first boundary trigger moment after energy redistribution. This is the weighting coefficient for energy transfer losses. This is the temperature rise weighting coefficient. The boundary improvement weighting coefficients are set as follows. In this embodiment, the energy transfer loss weighting coefficient is 1, the temperature rise weighting coefficient is 2, and the boundary improvement weighting coefficient is 0.1. The calculated transport costs are: -0.22 for the path from cabinet 5 to cabinet 3, -0.19 for the path from cabinet 1 to cabinet 4, -0.14 for the path from cabinet 2 to cabinet 4, and 0.24 for the path from cabinet 6 to cabinet 3. Therefore, the first three paths are selected for redistribution. The actual arrangement is: 5.2 kWh transferred from cabinet 5 to cabinet 3, 6.0 kWh transferred from cabinet 1 to cabinet 4, and 2.8 kWh transferred from cabinet 2 to cabinet 4, for a total transfer of 14 kWh. The conversion loss is 0.55 kWh, and the equivalent efficiency is 96.1%.

[0065] like Figure 4 As shown, Figure 4 This is a comparison chart of the initial boundary trigger times. The chart uses the initial boundary trigger times before and after energy redistribution as the vertical axis and the six battery cabinets as the horizontal axis, with the target boundary time of 120 minutes marked by a dashed line. As can be seen in the chart, before redistribution, cabinets 3 and 4 were both below the target boundary time; after redistribution, cabinet 3 rose to 135 minutes, and cabinet 4 rose to 132 minutes, improvements of 17 minutes and 20 minutes respectively, successfully surpassing the target boundary time. Cabinets 2 and 6 also improved to 145 minutes and 141 minutes respectively, indicating that the overall boundary distribution was more balanced after the scheduling recalculation. Therefore, this invention can not only identify the weakest battery cabinets but also bring the earliest lagging battery cabinets back into the acceptable range through local energy redistribution.

[0066] After energy redistribution is completed, the system continues to monitor the recovery response and updates cabinet-level capability status parameters, multi-boundary risk indicators, and cumulative aging costs accordingly. For example... Figure 5 As shown, Figure 5 This is a recovery response curve after energy redistribution. Taking cabinet 4 as an example, the horizontal axis represents recovery time, and the three curves represent the terminal voltage of cabinet 4, the maximum temperature of cabinet 4, and the boundary margin of cabinet 4, respectively. Within 20 minutes after energy redistribution, the terminal voltage of cabinet 4 recovered from 702.0 volts to 709.3 volts, the maximum temperature decreased from 36.8 degrees Celsius to 34.8 degrees Celsius, and the boundary margin recovered from 0.09 to 0.242. The predicted recovery results for cabinet 4 before redistribution were 708.7 volts, 35.1 degrees Celsius, and 0.23, respectively. Therefore, the differences between the actual recovery response and the predicted recovery response are 0.6 volts, 0.3 degrees Celsius, and 0.012, respectively, all within the preset allowable range. Based on this result, the system simultaneously increased the thermal safety margin and the individual voltage constraint margin of cabinet 4, and corrected the cumulative aging cost of cabinet 4 from 1.36 to 1.29, so that cabinet 4 is no longer overly conservatively restricted in the next scheduling cycle.

[0067] The overall implementation process demonstrates that, under the same target power task, this invention first identifies high-risk battery cabinets through multi-boundary risk assessment, then uses a reinforcement learning scheduler to allocate roles and power shares, and finally addresses local shortcomings through energy redistribution under optimal transport constraints. Results show that the initial boundary trigger time of the earliest lagging battery cabinet was improved from 112 minutes to 132 minutes, an improvement of 17.9%; the proportion of high-risk battery cabinets 3 and 4 in the total discharge was controlled at 19.4%; the energy redistribution efficiency reached 96.1%, and the recovery response remained consistent with the predicted recovery response. This indicates that the invention not only meets the station-level target power task but also suppresses premature boundary triggering and aging concentration during continuous scheduling, demonstrating clear engineering implementation value.

[0068] like Figure 6 As shown, a reinforcement learning-based battery pooling scheduling device is used to implement a reinforcement learning-based battery pooling scheduling method. The device includes: The status determination module is used to acquire operational data and disturbance response data from each battery cabinet, determining the cabinet-level capability status parameters of each cabinet. This module can be composed of a cabinet-level acquisition unit, a cabinet-level controller, and a station-level data interface. The cabinet-level acquisition unit connects to voltage sensors, current sensors, temperature sensors, insulation detection circuits, and sampling loops related to disturbance testing to acquire operational and disturbance response data from each battery cabinet. The cabinet-level controller filters, aligns the time, and calculates the status of the acquired signals, outputting the cabinet-level capability status parameters. This module is designed for placement closer to the battery cabinets and is part of the field sensing and front-end status calculation hardware.

[0069] The risk determination module is used to determine the multi-boundary risk indicators, cumulative aging costs, and initial boundary trigger times for each battery cabinet based on the capacity status parameters of each cabinet. This module can be implemented by the computing and storage units within a station-level industrial controller or energy storage management controller. The computing unit receives the cabinet-level capacity status parameters uploaded by the status determination module and performs calculations of multi-boundary risk indicators, updates of cumulative aging costs, and predictions of the initial boundary trigger times. The storage unit is used to save historical operating states, boundary thresholds, aging records, and intermediate calculation results. This module is more geared towards centralized computing hardware and primarily undertakes risk analysis and state evolution judgment functions.

[0070] The scheduling generation module determines the power allocation and recovery allocation parameters for each battery cabinet based on various multi-boundary risk indicators, cumulative aging costs, the initial boundary trigger time, and the target power task. Based on these parameters, it uses a reinforcement learning scheduler to generate the scheduling role and power share for each battery cabinet. The generated results are then subjected to operational constraint verification to obtain pooled scheduling instructions. This module can be implemented on a station-level main control computing platform, using an industrial computer, edge server, or controller with a dedicated acceleration unit. It receives the multi-boundary risk indicators, cumulative aging costs, and the initial boundary trigger time from the risk determination module, and combines these with the target power task to generate the power allocation parameters, recovery allocation parameters, and scheduling roles. Operational constraint verification is also performed within this module. The generated pooled scheduling instructions are then sent to each cabinet-level controller and power conversion device via a communication interface. This module is part of the global scheduling decision hardware.

[0071] The energy redistribution module, responding to situations where the initial boundary trigger time after the execution of the pooled scheduling instruction is earlier than the target boundary time corresponding to the target power task, determines the energy redistribution path based on the optimal transport method and executes energy redistribution. It updates the cabinet-level capacity status parameters, multi-boundary risk indicators, and cumulative aging costs based on the recovery response after energy redistribution, and outputs the pooled scheduling instruction for the next scheduling cycle. The redistribution module can be composed of a path solving unit in the station-level controller and a field execution unit. The path solving unit calculates the energy redistribution path based on the pooled scheduling instruction execution result and the target boundary time; the field execution unit drives the corresponding power converters, DC-DC converter interfaces, contactors, or other controlled switching components to execute energy redistribution. After execution, the sampling loop transmits recovery response data back to update the cabinet-level capacity status parameters, multi-boundary risk indicators, and cumulative aging costs. This module combines the characteristics of both control calculation hardware and power execution hardware.

[0072] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0073] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A battery pooling scheduling method based on reinforcement learning, characterized in that, The method includes: Obtain the operation data and perturbation response data of each battery cabinet to determine the cabinet-level capability status parameters of each battery cabinet; Based on the cabinet-level capability status parameters, determine the multi-boundary risk indicators, cumulative aging costs, and first boundary trigger time for each battery cabinet; Based on the multi-boundary risk indicators, the cumulative aging cost, the first boundary trigger time, and the target power task, the power allocation parameters and recovery allocation parameters of each battery cabinet are determined. Based on the power allocation parameters and the recovery allocation parameters, the scheduling role and power share of each battery cabinet are generated using a reinforcement learning scheduler. The generated results are then subjected to operational constraint verification to obtain pooled scheduling instructions. In response to the first boundary trigger time being earlier than the target boundary time corresponding to the target power task after the execution of the pooled scheduling instruction, the energy redistribution path is determined based on the optimal transport method and energy redistribution is performed. The cabinet-level capability status parameters, the multi-boundary risk indicators and the cumulative aging cost are updated according to the recovery response after energy redistribution, and the pooled scheduling instruction for the next scheduling cycle is output.

2. The method according to claim 1, characterized in that, The process of acquiring operational data and perturbation response data for each battery cabinet and determining the cabinet-level capability status parameters for each battery cabinet includes: Obtain the operating data of each battery cabinet, including the total voltage of the battery cabinet, the total current of the battery cabinet, the highest voltage of each individual cell, the lowest voltage of each individual cell, temperature data, and insulation status; Acquire perturbation response data for each battery cabinet, including instantaneous voltage drop data, recovery slope data, and temperature response data; Based on the operational data and the perturbation response data, the cabinet-level capability status parameters of each battery cabinet are determined.

3. The method according to claim 2, characterized in that, The cabinet-level capability status parameters include release energy margin, absorbable energy margin, thermal safety margin, individual unit voltage constraint margin, and interface stability margin.

4. The method according to claim 1, characterized in that, The determination of the multi-boundary risk indicators and the first boundary trigger time for each battery cabinet based on the cabinet-level capability status parameters includes: Based on the cabinet-level capability status parameters, the charging boundary risk index, discharging boundary risk index, temperature boundary risk index, single-cell voltage boundary risk index, and interface stability boundary risk index are determined respectively. Based on the charging boundary risk index, the discharging boundary risk index, the temperature boundary risk index, the single cell voltage boundary risk index, and the interface stability boundary risk index, the boundary margin sequence of each battery cabinet is determined. The first moment in the boundary margin sequence that falls below the preset boundary threshold is determined as the first boundary trigger moment for each battery cabinet.

5. The method according to claim 4, characterized in that, The determination of the cumulative aging cost of each battery cabinet includes: Based on the power load of each battery cabinet in the current scheduling cycle and its historical operating status, determine the aging increment of each battery cabinet; The amount of reduction in aging cost for each battery cabinet is determined based on the amount of boundary margin recovery for each battery cabinet during the recovery phase. The cumulative aging cost of each battery cabinet is updated based on the aging increment and the aging cost reduction.

6. The method according to claim 1, characterized in that, The process of determining the power allocation parameters and recovery allocation parameters for each battery cabinet based on the multi-boundary risk indicators, the cumulative aging cost, the initial boundary trigger time, and the target power task includes: Based on the aforementioned multi-boundary risk indicators, the cumulative aging cost, and the first boundary trigger time, a set of candidate roles for each battery cabinet is generated. The set of candidate roles includes a set of candidate roles for discharging, charging, buffering, recovery, observation, and isolation. Based on the candidate sets of each role and the target power task, determine the power allocation parameters and recovery allocation parameters for each battery cabinet.

7. The method according to claim 6, characterized in that, Based on the power allocation parameters and the recovery allocation parameters, a reinforcement learning scheduler is used to generate the scheduling role and power share of each battery cabinet, and a pooled scheduling instruction is obtained under the condition of satisfying the operating constraints, including: The power allocation parameters, the recovery allocation parameters, the role candidate set, and the target power task are input into the reinforcement learning scheduler to generate candidate scheduling schemes. Based on the initial boundary trigger time and the cumulative aging cost, the candidate scheduling schemes are evaluated to determine the target scheduling scheme. Perform operational constraint verification on the target scheduling scheme; If the operation constraint verification fails, the target scheduling scheme is modified to obtain the pooling scheduling instruction, wherein the operation constraints include individual cell voltage constraints, battery cabinet current constraints, temperature constraints, insulation constraints, and contactor constraints.

8. The method according to claim 7, characterized in that, The process of determining energy redistribution paths and performing energy redistribution based on optimal transport methods includes: The source battery cabinet and the target battery cabinet are determined based on the deviation between the initial boundary trigger time and the target boundary time corresponding to the target power task. The transport cost is determined based on energy transfer loss, temperature rise, and the increase at the first boundary trigger moment after energy redistribution. The energy redistribution path is determined based on the transportation cost, and the energy redistribution corresponding to the energy redistribution path is executed when the load fluctuation is lower than a preset threshold.

9. The method according to claim 8, characterized in that, The step of updating the cabinet-level capability status parameters, the multi-boundary risk indicators, and the cumulative aging cost based on the recovery response after energy redistribution includes: Collect the recovery response after energy redistribution to determine the margin recovery amount; Based on the boundary margin recovery amount, update the cabinet-level capability status parameters, the multi-boundary risk indicators, and the cumulative aging cost; The recovery response after energy redistribution is compared with the predicted recovery response before energy redistribution. If the difference obtained from the comparison exceeds a preset threshold, the power allocation parameter and the recovery allocation parameter are corrected.

10. A battery pooling scheduling apparatus based on reinforcement learning, used to implement the battery pooling scheduling method based on reinforcement learning as described in any one of claims 1-9, characterized in that, The device includes: The status determination module is used to acquire the operating data and perturbation response data of each battery cabinet and determine the cabinet-level capability status parameters of each battery cabinet. The risk determination module is used to determine the multi-boundary risk indicators, cumulative aging costs, and first boundary trigger time of each battery cabinet based on the cabinet-level capability status parameters. The scheduling generation module is used to determine the power allocation parameters and recovery allocation parameters of each battery cabinet based on the multi-boundary risk indicators, the cumulative aging cost, the first boundary trigger time, and the target power task. Based on the power allocation parameters and the recovery allocation parameters, it uses a reinforcement learning scheduler to generate the scheduling role and power share of each battery cabinet, performs runtime constraint verification on the generation results, and obtains pooled scheduling instructions. The redistribution module is used to respond to the first boundary trigger time being earlier than the target boundary time corresponding to the target power task after the execution of the pooled scheduling instruction, determine the energy redistribution path based on the optimal transport method and perform energy redistribution, update the cabinet-level capacity status parameters, the multi-boundary risk indicators and the cumulative aging cost according to the recovery response after energy redistribution, and output the pooled scheduling instruction for the next scheduling cycle.