Energy scheduling system and method for an energy storage system

By employing polynomial chaotic expansion and deep deterministic strategy algorithms, a risk framework and decision-feasibility domain for energy storage systems are constructed. This solves the problem of uncertainty quantification in energy scheduling of energy storage systems, realizes a high-precision and highly adaptable scheduling strategy, and ensures the safe and stable operation of the system.

CN122118969APending Publication Date: 2026-05-29HUNAN XILAIKE ENERGY STORAGE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN XILAIKE ENERGY STORAGE TECH CO LTD
Filing Date
2026-04-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing energy storage system energy dispatching technologies cannot effectively quantify the uncertainty of operating status, resulting in insufficient compliance of dispatching strategies, easy failure of strategies and increased system operation risks, and lack of closed-loop adaptive capabilities.

Method used

The probability distribution of state data is quantified by multinomial chaos expansion, a risk system and decision-feasible domain are constructed, the objective function is solved by deep deterministic strategy algorithm, and trajectory verification and dynamic adjustment are performed by pre-simulation module and feedback module to form a closed-loop optimization mechanism.

Benefits of technology

It improves the accuracy and adaptability of energy dispatching in energy storage systems, ensures the practical feasibility and safety of dispatching strategies, dynamically adapts to changes in operating conditions, reduces operational risks, and enhances robustness and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122118969A_ABST
    Figure CN122118969A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of energy storage scheduling, and provides an energy scheduling system and method of an energy storage system, which quantifies a probability distribution of state data through a fusion module, constructs a risk system based on the probability distribution, and generates a decision feasible region to provide a constraint basis for a scheduling strategy; a scheduling module constructs a target function by using an uncertainty boundary of the decision feasible region and a weighted expected value of the risk system, solves the target function through a deep deterministic policy algorithm, and outputs an initial scheduling strategy; a pre-play module realizes pre-play of an execution track of the initial scheduling strategy within a time window, compares the execution track with a safety boundary, and ensures actual feasibility and safety of the scheduling instruction; and a feedback module analyzes a time sequence deviation between an actual response track and the execution track, dynamically adjusts the target function and the safety boundary based on statistical characteristics of the time sequence deviation, enables the scheduling system to dynamically adapt to changes in actual operation conditions, and eliminates operation risks caused by track deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of energy storage dispatching technology, and in particular to an energy dispatching system and method for an energy storage system. Background Technology

[0002] Energy storage systems are supporting equipment for new energy grid integration, peak shaving and valley filling, and stable operation of the power grid. Energy dispatch strategies directly determine the operating efficiency, safety, stability, and economy of energy storage systems. With the increasing randomness of new energy output and the increasing characteristics of load fluctuations, energy dispatch technology for energy storage systems is gradually evolving from traditional deterministic dispatch modes to dynamic dispatch modes that consider operational uncertainties.

[0003] Existing technologies mostly collect operational status data of energy storage systems, combine them with conventional probabilistic and statistical methods to characterize operational uncertainties, and use heuristic algorithms or traditional optimization algorithms to construct scheduling objective functions and solve strategies. At the same time, they rely on basic feedback loops to realize simple adjustments to scheduling parameters, which improves the automation level of energy storage system scheduling to a certain extent and adapts to energy scheduling needs under normal operating conditions.

[0004] Existing energy storage system energy dispatching technologies still have many shortcomings. Existing technologies cannot achieve sparsity and uncertainty representation of state data, and are out of touch with actual operating characteristics; they can only achieve simple judgment of execution trajectory, which is prone to insufficient policy compliance; existing dispatching systems do not have closed-loop adaptive capabilities, and the deviation between actual response and pre-simulated trajectory will continue to accumulate, leading to dispatching strategy failure and increased system operation risk.

[0005] Based on the shortcomings of the existing technology, the technical problem to be solved in this application is how to quantify the uncertainty of the operating state of the energy storage system in order to achieve the safety and sustainability of energy dispatch. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this application provides an energy dispatching system and method for an energy storage system.

[0007] In a first aspect, this application provides an energy dispatching system for an energy storage system, the system comprising: a fusion module, a dispatching module, a pre-simulation module, and a feedback module;

[0008] The fusion module is used to acquire state data reflecting the operating status of the energy storage system, quantifies the probability distribution of the state data through multinomial chaos expansion, constructs a risk system based on the probability distribution, and generates a decision feasible domain including uncertainty boundaries and risk weights.

[0009] The scheduling module is used to construct an objective function by weighting the expected value of the uncertainty boundary and risk system corresponding to the decision feasible domain, solve the objective function through a deep deterministic strategy algorithm, and output the initial scheduling strategy.

[0010] The pre-execution module is used to receive and pre-execute the execution trajectory of the initial scheduling strategy in different time windows, and determine whether the execution trajectory is within the safety boundary. If so, the initial scheduling strategy is output as a scheduling instruction. If not, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation until the scheduling instruction is output.

[0011] The feedback module is used to obtain the actual response trajectory of the scheduling instruction and perform time-series deviation analysis with the execution trajectory. Based on the statistical characteristics of the time-series deviation, the objective function and safety boundary are dynamically adjusted.

[0012] As an optional implementation, the probability distribution of the quantized state data includes:

[0013] Obtain state data reflecting the operating status of the energy storage system, calculate the contribution of the state data to the uncertainty of the energy storage system operation using the mutual information method, filter out state data with a contribution greater than the contribution threshold, and form coupled states through time series correlation.

[0014] The mean square error between the empirical distribution and the standard random distribution of the coupled state is calculated. The basis functions are matched, and the deceleration rate of the chaotic expansion error is used as the convergence exponent until it is less than the convergence threshold. Then the expansion order is output.

[0015] The coupled state is expanded using a polynomial chaotic expansion with basis functions and expansion order. The expansion coefficients are solved by minimum angle regression. After removing terms whose absolute values ​​of the expansion coefficients are less than a preset threshold, a sparse chaotic expansion model is obtained.

[0016] The probability distribution, including probability density and cumulative probability, in the coupled state is reconstructed based on the sparse chaotic expansion model, and the quantile corresponding to the preset confidence level is used as the uncertainty boundary.

[0017] As an optional implementation, generating the decision feasible domain includes:

[0018] Based on the probability distribution including probability density and cumulative probability in the coupled state, a risk system including local risk and global risk is constructed by dividing the scope of influence.

[0019] The volatility coefficient of the coupled state is determined based on the variance and covariance of the probability distribution. Risk weights are allocated according to the magnitude of the volatility coefficients, and the risk weights are adjusted according to the differences in risk levels in the risk system.

[0020] Based on uncertainty boundaries and risk weights, and combined with the collaborative constraints formed by the temporal correlation between coupled states, a decision-feasible region is generated.

[0021] As an optional implementation, the objective function includes:

[0022] The local and global risk values ​​are normalized, and the priority of the indicators for local and global risks is determined by combining the magnitude of the volatility coefficient of the coupled state.

[0023] The weighting coefficients are obtained by calibrating the initial coefficients using the corrected risk weights and combining them with the variance and covariance of the probability distribution. The weighting coefficients are then adjusted according to the priority of the indicators.

[0024] The normalized local risk value, the normalized global risk value, and the corresponding modified weighting coefficient are multiplied to obtain the weighted expected values ​​of local risk and global risk, respectively. The coordination factor is determined according to the magnitude of the cumulative probability, and the weighted expected value of the risk system is obtained by weighted fusion.

[0025] Based on the weighted expected value of the risk system, and combined with the constraint penalty terms set by the uncertainty boundary and collaborative constraints corresponding to the feasible decision domain, the objective function is constructed.

[0026] As an optional implementation, the output initial scheduling strategy includes:

[0027] By combining indicator priority, fluctuation coefficient of coupled state and collaborative constraints, a dual network for deep deterministic strategy algorithm is initialized;

[0028] The samples generated by the interaction between the deep deterministic strategy algorithm and the objective function are obtained, and the samples are stored in a hierarchical manner based on the probability density and the difference in risk level in the risk system. Samples are then sampled from the experience replay pool.

[0029] The dual network is constrained and calibrated based on the uncertainty boundary and cumulative probability. The objective function is then iteratively solved under the constraint calibration. The iteration stops when the solution error of the objective function, the triggering frequency of the constraint penalty term, and the termination condition of the global risk value are all satisfied. The solution result is then output as the initial scheduling policy.

[0030] As an optional implementation, the safety boundary is characterized as an operational safety constraint interval composed of the probability distribution of the coupled states and the local and global risk values. The operational safety constraint interval is based on the uncertainty boundary and uses the temporal correlation between the coupled states as a collaborative constraint, which is used to compare the interval affiliation with the execution trajectory.

[0031] If the execution trajectory is within the safety boundary, the initial scheduling policy will be output as a scheduling instruction.

[0032] If the execution trajectory is not within the safety boundary, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation.

[0033] As an optional implementation, the reverse correction uncertainty boundary includes:

[0034] The system receives the initial scheduling policy, divides time windows based on the temporal correlation of the coupling state, performs temporal deduction on the changes in the coupling state of the initial scheduling policy in different time windows, and obtains the corresponding execution trajectory.

[0035] If the execution trajectory is not located within the safety boundary, the degree of trajectory deviation is determined based on the spatial deviation of the execution trajectory relative to the safety boundary, the risk deviation value, and the duration of the temporal deviation.

[0036] The quantiles corresponding to the uncertainty boundary are corrected in reverse according to the degree of trajectory deviation, wherein the correction direction of the uncertainty boundary is opposite to the direction of trajectory deviation, and the correction amount is positively correlated with the degree of trajectory deviation.

[0037] After the correction is completed, the timing of the initial scheduling strategy is re-performed and the execution trajectory is compared until the execution trajectory is within the safety boundary, and then the scheduling instruction is output.

[0038] As an optional implementation, the dynamically adjusted objective function includes:

[0039] Obtain the actual response trajectory of the scheduling instruction, perform timing deviation analysis between the actual response trajectory and the execution trajectory obtained from the pre-simulation, and extract the statistical characteristics of the timing deviation, including amplitude, frequency and cumulative amount.

[0040] Based on the statistical characteristics of time series deviations, the weighting coefficients, coordination factors, and constraint penalty terms in the objective function are dynamically adjusted.

[0041] The execution trajectory of the initial scheduling strategy under the constraints of the adjusted objective function is simulated, and the statistical characteristics of the timing deviation after adjustment are verified until the verification is passed to determine the parameter configuration of the objective function.

[0042] As an optional implementation, dynamically adjusting the security boundary includes:

[0043] Based on the statistical characteristics of the time-series deviation, and combined with the probability distribution, local risk value, and global risk value of the coupling state corresponding to the safety boundary, the interval threshold and cooperative constraints of the safety boundary are dynamically adjusted.

[0044] The adjusted safety boundary and uncertainty boundary are subjected to constraint consistency verification, the execution trajectory of the initial scheduling strategy is rehearsed, and the interval attribution comparison results are verified until the verification is passed to determine the safety boundary.

[0045] Secondly, this application provides an energy dispatching method for an energy storage system. The method includes: acquiring state data reflecting the operating state of the energy storage system, quantifying the probability distribution of the state data through multinomial chaos expansion, constructing a risk system based on the probability distribution, and generating a decision feasible region including uncertainty boundaries and risk weights.

[0046] The objective function is constructed by the weighted expected value of the uncertainty boundary and risk system corresponding to the decision feasible region. The objective function is solved by the deep deterministic strategy algorithm to output the initial scheduling strategy.

[0047] Receive and rehearse the execution trajectory of the initial scheduling strategy within different time windows, and determine whether the execution trajectory is within the safety boundary. If so, output the initial scheduling strategy as a scheduling instruction. If not, reverse the uncertainty boundary according to the degree of trajectory deviation until the scheduling instruction is output.

[0048] The actual response trajectory of the scheduling instruction is obtained, and a timing deviation analysis is performed between the trajectory and the execution trajectory. Based on the statistical characteristics of the timing deviation, the objective function and safety boundary are dynamically adjusted.

[0049] Compared with the prior art, the beneficial effects of this application are: the fusion module quantifies the probability distribution of state data through multinomial chaos expansion, the risk system built based on the probability distribution is more in line with the actual operating characteristics of the system, and the generated decision feasible domain synchronously integrates uncertainty boundary and risk weight, providing a scientific and reasonable constraint basis for scheduling strategy.

[0050] The scheduling module constructs an objective function based on the uncertainty boundary of the decision-feasible domain and the weighted expected value of the risk system. The objective function is highly matched with the system's risk control and safe operation requirements. The objective function is solved through a deep deterministic strategy algorithm, which can adapt to the action space of continuous scheduling of energy storage systems. The output initial scheduling strategy has higher accuracy and adaptability.

[0051] The pre-simulation module can simulate the execution trajectory of the initial scheduling strategy within multiple time windows. By determining the attribution of the trajectory to the safety boundary and combining the degree of trajectory deviation, the uncertainty boundary can be corrected in reverse. This can avoid the risk of strategy execution in advance and ensure that the output scheduling instructions have practical feasibility and security.

[0052] The feedback module analyzes the time-series deviation between the actual response trajectory and the pre-simulated execution trajectory, and uses the statistical characteristics of the time-series deviation to dynamically adjust the objective function and safety boundary, thus constructing a scheduling closed-loop optimization mechanism. This enables the scheduling system to dynamically adapt to changes in actual operating conditions, eliminate operational hazards caused by trajectory deviations, and improve the robustness, stability, and operational efficiency of energy scheduling in the energy storage system.

[0053] This application achieves uncertainty quantification, strategy optimization, trajectory verification, and closed-loop adjustment of energy scheduling in energy storage systems through the coordinated setup of fusion module, scheduling module, pre-simulation module, and feedback module. It effectively solves the inherent defects of existing scheduling technologies and takes into account the safety constraints, risk management, and efficient operation requirements of energy scheduling in energy storage systems. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0055] Figure 1 A system flowchart of an energy dispatching system for an energy storage system provided in an embodiment of this application;

[0056] Figure 2 A logical flowchart illustrating the output initial scheduling strategy provided in the embodiments of this application;

[0057] Figure 3 This is a flowchart illustrating an energy scheduling method for an energy storage system provided in an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0059] Example 1:

[0060] like Figure 1 The diagram shown is a system flowchart of an energy dispatching system for an energy storage system provided in this application embodiment. The system includes a fusion module, a dispatching module, a pre-simulation module, and a feedback module.

[0061] The fusion module is used to acquire state data reflecting the operating status of the energy storage system, quantifies the probability distribution of the state data through multinomial chaos expansion, constructs a risk system based on the probability distribution, and generates a decision feasible domain including uncertainty boundaries and risk weights.

[0062] Furthermore, the probability distribution of the quantified state data includes:

[0063] Obtain state data reflecting the operating status of the energy storage system, calculate the contribution of the state data to the uncertainty of the energy storage system operation using the mutual information method, filter out state data with a contribution greater than the contribution threshold, and form coupled states through time series correlation.

[0064] The mean square error between the empirical distribution and the standard random distribution of the coupled state is calculated. The basis functions are matched, and the deceleration rate of the chaotic expansion error is used as the convergence exponent until it is less than the convergence threshold. Then the expansion order is output.

[0065] The coupled state is expanded using a polynomial chaotic expansion with basis functions and expansion order. The expansion coefficients are solved by minimum angle regression. After removing terms whose absolute values ​​of the expansion coefficients are less than a preset threshold, a sparse chaotic expansion model is obtained.

[0066] The probability distribution, including probability density and cumulative probability, in the coupled state is reconstructed based on the sparse chaotic expansion model, and the quantile corresponding to the preset confidence level is used as the uncertainty boundary.

[0067] The operational status data of energy storage systems is complex and contains a high proportion of redundant information. Directly performing probability distribution quantification would significantly increase computational overhead, and low-correlation parameters would interfere with the accuracy of uncertainty characterization. Therefore, we need to acquire status data that reflects the operational status of energy storage systems, including the state of charge of energy storage units, charging and discharging power, DC bus voltage, ambient temperature and grid frequency fluctuations, etc., to form quantitative indicators that characterize the system's operational characteristics. We then use the mutual information method to calculate the contribution of status data to the operational uncertainty of energy storage systems.

[0068] Using the state data as input variables X and the comprehensive characterization of system operation uncertainty as output variables Y, the mutual information value I(X;Y) is calculated through a discrete probability distribution. The resulting mutual information value is the contribution of the corresponding state data. The larger the value, the more significant the influence of the parameter on the system uncertainty.

[0069] Filtering state data with a contribution greater than the contribution threshold means removing redundant parameters with a contribution lower than the contribution threshold through numerical comparison. The contribution threshold is determined based on the median level of the contribution of the energy storage system's conventional operating parameters and is used to filter parameters with weak influence. Coupled states are formed through time-series correlation, that is, the time-series correlation coefficient between parameters is calculated in fixed time windows, and parameters with a time-series correlation coefficient greater than the set value are integrated into coupled state variables to characterize the time-series coordinated change characteristics of multiple parameters.

[0070] For example, the state data includes State of Charge (SOC), charging and discharging power (P), DC bus voltage (U), ambient temperature (T), and grid frequency fluctuation (Δf). After calculation using the mutual information method, the contribution values ​​of each parameter are 0.32, 0.28, 0.21, 0.09, and 0.1, respectively. The contribution threshold is set to 0.15, which is the median critical value of the contribution value in this calculation. Weakly influential parameters are eliminated, and ambient temperature T with a contribution value of 0.09 is also eliminated, retaining the three types of parameters: SOC, P, and U. With a fixed time window of 5 minutes, the time-series correlation coefficient between P and SOC is calculated to be 0.89, which is higher than the coupling judgment value of 0.7. Therefore, the two are integrated into a coupled state variable {P, SOC}.

[0071] Redundant data is eliminated to reduce subsequent computational load, and the characteristics of system collaborative operation are constructed through temporal coupling to characterize variables, thus avoiding the one-sidedness of single-parameter analysis. The data purity and coupling rationality of the coupled state directly determine the computational accuracy of subsequent steps and are the basic premise of the entire quantization process.

[0072] The accuracy of polynomial chaotic expansion is determined by the basis function fit and the reasonableness of the expansion order. A mismatch in basis functions leads to a large expansion error, while an excessively high order easily results in overfitting, and an excessively low order provides insufficient representation. The mean square error between the empirical distribution and the standard random distribution of the coupled states is calculated to match the basis functions. Here, the empirical distribution is sampled from the empirical cumulative distribution function F. emp (x) represents the statistical distribution of the coupled state sampling data, characterizing the proportion of samples whose values ​​are not greater than variable x. The cumulative distribution function of the standard random distribution is G(x), and the mean squared error MSE = Σ[F emp (x i )-G(x i )]² / n, where n is the number of sampling points, x i Let MSE be the coupling state value at the i-th sampling point. The smaller the MSE value, the higher the distribution fit. The orthogonal basis function corresponding to the standard distribution with the smallest MSE is selected as the optimal basis function.

[0073] The deceleration rate of the chaotic expansion error is used as the convergence exponent. When the expansion order is less than the convergence threshold, the expansion order is output, and the chaotic expansion error E is calculated step by step. k (k is the expansion order), error deceleration rate r = (E k -E k+1 ) / E k The value r is the convergence exponent. The smaller the convergence exponent, the weaker the improvement in accuracy from increasing the order. When r is less than the convergence threshold, the current order is considered the optimal order. There are two types of convergence threshold settings. The first is the engineering critical value in the stochastic quantization scenario of energy storage system using polynomial chaotic expansion, which is usually in the range of 0.01 to 0.03. Within this range, the accuracy improvement brought by further increasing the order is negligible. The second is determined through pre-iteration experiments. For the current coupling state distribution, 1 to 5 orders are pre-expanded, and the inflection point value of the error deceleration rate is observed. The inflection point value is used as the convergence threshold. Here, the convergence threshold is set to 0.02. This value satisfies both the engineering critical value requirement and the pre-iteration inflection point characteristic. Further increasing the order has no substantial accuracy gain.

[0074] For example, for the empirical distribution of the coupled state {P, SOC}, the errors calculated using the above formula are 0.023, 0.071, and 0.056 respectively, compared with the normal distribution, uniform distribution, and gamma distribution. The normal distribution has the smallest MSE, so the Gaussian orthogonal basis function is selected as the optimal basis function. The convergence threshold is set to 0.02, and the expansion error is calculated step by step: 1st order E1 = 0.082, 2nd order E2 = 0.031, r1 = (0.082 - 0.031) / 0.082 ≈ 0.622; 2nd order E2 = 0.031, For the 3rd order E3=0.024, r2=(0.031-0.024) / 0.031≈0.226; for the 3rd order E3=0.024 and the 4th order E4=0.0235, r3=(0.024-0.0235) / 0.024≈0.0208; for the 4th order E4=0.0235 and the 5th order E5=0.0231, r4=(0.0235-0.0231) / 0.0235≈0.017; since r4=0.017<0.02, the 5th order is determined to be the optimal expansion order.

[0075] The basis functions are accurately matched by mean square error, and the order is determined by the rate of decrease of error, thus balancing the accuracy of representation and the computational efficiency. The determined basis functions and order provide computational parameters for subsequent chaotic expansion, which limits the mathematical framework of chaotic expansion and directly determines the foundation for the construction of sparse chaotic expansion models.

[0076] Conventional fixed-order polynomial chaotic expansions require pre-setting high-order expansion forms, and the number of expansion terms increases exponentially with the order. This results in a large amount of redundant computation during online real-time calculations, making it unsuitable for the real-time quantitative scheduling requirements of energy storage systems. A polynomial chaotic expansion of the coupled state is performed using the aforementioned matched optimal basis function and expansion order, representing the coupled state random variable as ξ=Σa. b ×Φ b (X), where a b Φ is the expansion coefficient corresponding to the b-th basis function. b (X) is the b-th basis function in the optimal basis function set. The order of expansion limits the highest order of the basis functions, thus realizing the linear decomposition representation of the random variable.

[0077] The expansion coefficients are solved by least angular regression, specifically aiming to minimize the fitting residual of the chaotic expansion. The coefficient vector is initialized as a zero vector, and the basis function terms with the highest correlation to the residual are selected successively. The coefficient vector is updated along the equiangular direction and iterated until the residual converges, finally obtaining the coefficient values ​​of each expansion term. This algorithm can automatically distinguish between core coefficients and redundant coefficients. After removing terms whose absolute values ​​of expansion coefficients are less than a preset threshold, a sparse chaotic expansion model is obtained. That is, small-coefficient redundant terms are removed by numerical comparison. The coefficient threshold is set according to the numerical distribution characteristics of the expansion coefficients, and core terms that make substantial contributions to the representation results are retained.

[0078] For example, based on Gaussian orthogonal basis functions and a 5th-order expansion, the above polynomial chaotic expansion is performed on the coupled state {P,SOC}. Through minimum angular regression iteration, the expansion coefficients are obtained as 0.42, 0.35, 0.11, 0.06, and 0.03 respectively. The preset threshold value for the coefficients is 0.08, which is based on the fact that this value is the critical value of the coefficient distribution in this case. Coefficients below this value contribute less than 5% to the expansion result and can be judged as redundant terms. The expansion terms corresponding to coefficients 0.06 and 0.03 are removed, and the core terms corresponding to coefficients 0.42, 0.35, and 0.11 are retained to obtain the sparse chaotic expansion model.

[0079] The efficiency of coefficient solving is improved by using minimum angular regression, and the sparsity of the model is achieved by eliminating redundant coefficients, thereby reducing the computational cost of subsequent probability distribution reconstruction. This sparse chaotic expansion model serves as the sole basis for subsequent probability distribution reconstruction, and its structure and accuracy directly determine the accuracy of the reconstruction of probability density and cumulative probability.

[0080] Sparse chaotic expansion models are abstract representations and cannot be applied to the construction of feasible decision regions and security constraints. This paper reconstructs the probability density and cumulative probability of coupled states based on sparse chaotic expansion models. The random variable expression of the coupled states is obtained through the sparse chaotic expansion model, and a probability integral transformation is performed on this expression. The cumulative probability distribution function adopts F... CDF (x)=P r (ξ≤x), where P r The probability operator, or probability mathematical notation, is used to represent the probability that the value of a random variable ξ is no greater than x, and its probability density function is f(x) = dF. CDF (x) / dx is obtained by numerical integration to obtain the probability density and cumulative probability values ​​for different values.

[0081] The quantile corresponding to the preset confidence level is used as the uncertainty boundary. That is, in the cumulative probability distribution, the value of the coupling state when the cumulative probability is equal to the preset confidence level is the quantile. This quantile is the uncertainty constraint boundary of the system operation. The confidence level is set according to the safety operation level of the energy storage system.

[0082] For example, based on the sparse chaotic expansion model, the probability density and cumulative probability distribution of the coupled state {P,SOC} are reconstructed through the above probability integral and numerical integral calculations; the confidence level is set at 95%, which is based on the confidence level of the industrial-grade safe operation standard for energy storage systems and can cover most conventional operation scenarios; in the cumulative probability distribution, when the cumulative probability reaches 95%, the SOC value is 0.85 and the P value is 0.9MW, and the parameter interval corresponding to this quantile is the uncertainty boundary.

[0083] The abstract mathematical model is transformed into an intuitive probability distribution, and the engineering constraint boundary is formed by confidence level quantiles, realizing the transformation of random characteristics into constraint parameters. The generated uncertainty boundary is directly applied to the subsequent generation of feasible domain and construction of objective function, which is the constraint basis of the entire energy storage scheduling system.

[0084] Specifically, the generated decision feasible domain includes:

[0085] Based on the probability distribution including probability density and cumulative probability in the coupled state, a risk system including local risk and global risk is constructed by dividing the scope of influence.

[0086] The volatility coefficient of the coupled state is determined based on the variance and covariance of the probability distribution. Risk weights are allocated according to the magnitude of the volatility coefficients, and the risk weights are adjusted according to the differences in risk levels in the risk system.

[0087] Based on uncertainty boundaries and risk weights, and combined with the collaborative constraints formed by the temporal correlation between coupled states, a decision-feasible region is generated.

[0088] The impact range and level of different parameter components on the energy storage system in the coupled state are fundamentally different. Using only a single-dimensional risk characterization method cannot distinguish between parameter-level local impact and system-level global impact, which will lead to a lack of specificity in subsequent risk management and constraint setting, and insufficient accuracy in risk assessment. Risk classification is based on the probability density and cumulative probability of the coupled state. That is, based on the probability density representing the occurrence probability of each component of the coupled state and the cumulative probability representing the overall distribution characteristics of the coupled state, the fluctuation impact range of each component is identified. The probability density is used to determine the abnormal fluctuation probability of a single component, and the cumulative probability is used to determine the system-level risk probability of coordinated fluctuation of multiple components.

[0089] Local risks are classified into local and global risks based on their scope of impact. Risks that only affect a single unit, parameter, or component of the energy storage system and do not cause overall operational anomalies are defined as local risks. The impact of such risks is limited to a single component of the coupled state. Risks that are caused by the coordinated fluctuations of multiple coupled state components and affect the overall scheduling logic, grid connection stability, and safe operation of the system are defined as global risks. The impact of such risks covers the entire energy storage system and is the key level to be controlled in the risk system.

[0090] For example, the coupled state is {P, SOC} consisting of charge / discharge power P and state of charge (SOC). Abnormal fluctuations in charge / discharge power P only correspond to abnormal output of a single energy storage battery cell, and the impact is limited to a local component. Based on the scope of impact, the risk corresponding to this fluctuation is determined to be a local risk. Coordinated fluctuations in SOC and charge / discharge power P directly affect the overall capacity allocation and grid interaction stability of the energy storage system, and the impact covers the entire system. Based on the scope of impact, the risk corresponding to this coordinated fluctuation is determined to be a global risk. Through probability distribution statistics, the probability of occurrence of local risk is 12%, and the probability of occurrence of global risk is 8%, thus completing the hierarchical classification of local and global risks.

[0091] This allows for the differentiation of control priorities between local and global risks, avoiding control biases caused by a single risk dimension and improving the pertinence and accuracy of risk characterization. The completed local and global risk system serves as the core basis for subsequent processing and determines the objectivity of subsequent weight settings.

[0092] The setting of risk weights needs to objectively reflect the actual volatility of the coupled state. At the same time, different risk levels have different impacts on the system. Uncorrected initial weights cannot match the hierarchical differences of the risk system, which will reduce the rationality of subsequent decision constraints. The volatility coefficient of the coupled state is determined based on the variance and covariance of the probability distribution. The variance is used to characterize the degree of discrete volatility of a single component of the coupled state, and the covariance is used to characterize the degree of linkage volatility between different components of the coupled state. The volatility coefficient is a quantitative index that combines the volatility of a single component and the linkage volatility between components. Specifically, the volatility coefficient is equal to the sum of the absolute values ​​of the component variance and the covariance between components divided by the nominal value of the component. After calculation, the volatility coefficient is normalized so that its value range is between 0 and 1, which is convenient for subsequent weight allocation.

[0093] Risk weights are allocated based on the magnitude of the volatility coefficient, following the principle that volatility coefficients and risk weights are positively correlated. The larger the volatility coefficient, the higher the volatility risk of that component, and the higher the initial risk weight allocated. The smaller the volatility coefficient, the lower the initial risk weight. Risk weights are adjusted according to the differences in risk levels within the risk system. Local risks are classified as low-risk levels, and global risks are classified as high-risk levels. A weight adjustment coefficient greater than 1 is set for high-risk levels, and a weight adjustment coefficient equal to 1 is set for low-risk levels. The product of the initial weight and the corresponding weight adjustment coefficient is used as the final adjusted risk weight. The weight adjustment coefficient is set according to the risk management level standard for energy storage systems.

[0094] For example, in the coupled state {P, SOC}, the variance of the charging / discharging power P is 0.02, and its nominal value is 1MW. The variance of the state of charge (SOC) is 0.03, and its nominal value is 1. The covariance between the two is 0.015. The calculated fluctuation coefficient of P is (0.02+0.015) / 1=0.035, and the fluctuation coefficient of SOC is (0.03+0.015) / 1=0.045. After normalization, the fluctuation coefficients are 0.44 and 0.56, respectively. Based on this, the initial risk weights P and SOC are assigned as 0.4 and 0.6, respectively. P corresponds to a local low-risk level, and the correction coefficient is 1. SOC and P together correspond to a global high-risk level, and the correction coefficient is 1.2. The corrected risk weights P and SOC are 0.72. After further normalization, the final risk weights P and SOC are 0.36 and 0.64, respectively.

[0095] This objectively quantifies the fluctuation characteristics of the coupling state, avoids the subjective arbitrariness of weight allocation, and completes weight correction in combination with risk level, so that the risk weight matches the actual risk impact, improves the rationality of subsequent decision constraints, and provides a risk dimension constraint basis for the final feasible domain generation.

[0096] Uncertainty boundaries limit the operational range of coupled states, risk weights limit the priority of risk management, and temporal coordination constraints between coupled states limit the dynamic correlation between parameters. A single constraint cannot form a complete scheduling decision space. Based on uncertainty boundaries and risk weights, the determined uncertainty boundaries are used as spatial range constraints of the decision-feasible domain, limiting the upper and lower limits of the values ​​of each component of the coupled state. At the same time, the modified risk weights are used as priority constraints of the decision-feasible domain, and more stringent constraint management is implemented on parameter components with high risk weights within the feasible domain.

[0097] The collaborative constraints formed by the temporal correlation between coupled states refer to the rules of linkage change of each component of the coupled state in the temporal dimension as dynamic constraints, which ensure that any value in the decision feasible domain satisfies the temporal collaborative relationship between parameters and that there will be no disconnection in parameter changes. The decision feasible domain is generated by superimposing and fusing spatial constraints, priority constraints, and dynamic collaborative constraints, forming a multi-dimensional constraint domain that covers all dimensions of the coupled state and takes into account both static range and dynamic correlation. All parameter combinations in this constraint domain are legal scheduling decisions that can be executed by the energy storage system.

[0098] For example, the uncertainty boundary is SOC≤0.85 and P≤0.9MW, the risk weight constraint is that the control priority of SOC is higher than that of P, and the time-series coordination constraint is that the ratio of the change in charge and discharge power P to the change in state of charge SOC must be between 0.8 and 1.2. By integrating the above three types of constraints and eliminating parameter combinations that do not meet the coordination constraints, a multi-dimensional decision-making feasible region that simultaneously meets the uncertainty boundary, risk weight priority, and time-series coordination rules is finally formed.

[0099] This constructs a decision-feasible domain that balances operational safety, risk control, and dynamic collaboration, addressing the issues of unreasonable decision space and easy failure of scheduling strategies under single constraints. It provides a legitimate and suitable decision range for the generation of subsequent scheduling strategies. The decision-feasible domain is the basis for subsequent deep deterministic strategy algorithms to solve the initial scheduling strategy, and determines the safety and feasibility of subsequent scheduling strategies.

[0100] The scheduling module is used to construct an objective function by weighting the expected value of the uncertainty boundary and risk system corresponding to the decision feasible domain, solve the objective function through a deep deterministic strategy algorithm, and output the initial scheduling strategy.

[0101] Furthermore, the objective function comprises:

[0102] The local and global risk values ​​are normalized, and the priority of the indicators for local and global risks is determined by combining the magnitude of the volatility coefficient of the coupled state.

[0103] The weighting coefficients are obtained by calibrating the initial coefficients using the corrected risk weights and combining them with the variance and covariance of the probability distribution. The weighting coefficients are then adjusted according to the priority of the indicators.

[0104] The normalized local risk value, the normalized global risk value, and the corresponding modified weighting coefficient are multiplied to obtain the weighted expected values ​​of local risk and global risk, respectively. The coordination factor is determined according to the magnitude of the cumulative probability, and the weighted expected value of the risk system is obtained by weighted fusion.

[0105] Based on the weighted expected value of the risk system, and combined with the constraint penalty terms set by the uncertainty boundary and collaborative constraints corresponding to the feasible decision domain, the objective function is constructed.

[0106] The local risk value and the global risk value differ in their dimensions and magnitudes, making it impossible to achieve a fair comparison and fusion calculation of the two types of risks. At the same time, the importance of local risks and global risks to the control of system scheduling differs. Normalizing the local risk value and the global risk value means using a minimum-maximum normalization method to eliminate the difference in dimensions and magnitudes between the two types of risk values, mapping the risk value to a unified numerical range of 0 to 1. Specifically, the original risk value is subtracted from the minimum value of the risk type, and then divided by the difference between the maximum and minimum values ​​of the risk type, so that the normalized risk value can directly participate in the weighted calculation.

[0107] The priority of indicators for local and global risks is determined by combining the magnitude of the volatility coefficient of the coupled state. This means that the volatility coefficient is used as a quantitative basis for the importance of risk management. The larger the volatility coefficient, the more significant the impact of the risk on the system operation, and the higher the priority of the corresponding indicator. Among them, the global risk is related to the coordinated fluctuation of multiple components in the coupled state. Its volatility coefficient comprehensively reflects the multi-parameter linkage characteristics, and its priority is naturally higher than that of local risks that are only related to a single component. The priority is clearly distinguished by hierarchical labels or sorting coefficients.

[0108] For example, local risk is the risk corresponding to the charging and discharging power P, with an original local risk value of 0.12. Global risk is the risk corresponding to the coordinated fluctuation of {P, SOC}, with an original global risk value of 0.08. The minimum value of the two types of risk values ​​is 0.08, and the maximum value is 0.12. After normalization, the local risk value is 1, and the global risk value is 0. The fluctuation coefficient of the charging and discharging power P in the coupled state is 0.44, and the global risk fluctuation coefficient involving the state of charge SOC is 0.56. Based on the magnitude of the fluctuation coefficient, the priority of the global risk indicator is determined to be higher than that of the local risk. The global risk is set as the first priority, and the local risk is set as the second priority.

[0109] This eliminates the differences in the dimensions and magnitudes of risk values, ensuring the fairness and rationality of subsequent calculations. By objectively classifying the priority of indicators through volatility coefficients, it avoids the subjective arbitrariness of priority setting, improves the scientific nature of risk management logic, and provides input parameters for the calibration and correction of subsequent weighting coefficients. The rationality of its values ​​and the accuracy of priority classification affect the suitability of the weighting coefficients.

[0110] The revised risk weights only reflect the basic correlation between risk level and volatility. Using them as the final coefficients would lead to a disconnect between the weighting logic and the actual operating characteristics of the system. At the same time, the priority of indicators reflects the importance of risk control. Using the revised risk weights as initial coefficients means directly using the final risk weights, which have been revised by the risk level, as the initial values ​​of the weighting coefficients. These initial values ​​provide the basic numerical basis for the weighting coefficients. The initial coefficients are calibrated by combining the variance and covariance of the probability distribution. The variance represents the degree of discrete volatility of the coupled state components, and the covariance represents the degree of linkage volatility between components. The calibration is completed by multiplying the initial coefficients by the weighted sum of the variance and covariance. The calibration process can eliminate the deviation between the initial coefficients and the actual probability distribution characteristics. The calibrated coefficients are more in line with the random volatility characteristics of the system.

[0111] The weighting coefficient is adjusted according to the priority of the indicators. This means that a corresponding priority adjustment coefficient is set for risks of different priorities. The adjustment coefficient for first-priority risks is greater than that for second-priority risks. The calibrated coefficient is multiplied by the corresponding priority adjustment coefficient to obtain the final adjusted weighting coefficient. The priority adjustment coefficient is set according to the risk management level standard of energy storage system.

[0112] For example, the corrected risk weights are 0.36 for charging / discharging power P and 0.64 for global risk, which are used as initial coefficients. The variance of charging / discharging power P is 0.02, and the covariance of the global risk correlation component is 0.015. After variance and covariance calibration, the local risk weighting coefficient is 0.35, and the global risk weighting coefficient is 0.65. Global risk is a first-level priority, with a correction coefficient of 1.1, and local risk is a second-level priority, with a correction coefficient of 1. The final corrected local risk weighting coefficient is 0.35, and the global risk weighting coefficient is 0.72.

[0113] This eliminates the mismatch between the initial coefficients and the probability distribution, and the weighting coefficients are matched with the importance of risk control by adjusting the priority of indicators, thereby improving the accuracy and rationality of the weighted calculation. The corrected weighting coefficients provide core operation parameters for the subsequent calculation of the expected value of local and global risk weights, and are the key intermediate variables for realizing the weighted fusion of risk values.

[0114] Local risk and global risk are independent risk dimensions, and the weighted expected value of a single dimension cannot characterize the overall risk level of the system. The weighted expected value is obtained by multiplying the normalized local risk value, the normalized global risk value, and the corresponding modified weighting coefficient. This means that the normalized risk value is used as the risk contribution base, and multiplied by the corresponding modified weighting coefficient. The product is the weighted expected value of this type of risk. This value can quantify the actual impact of a single type of risk after weighting.

[0115] The coordination factor is determined based on the magnitude of the cumulative probability. The cumulative probability reflects the probability that the coupling state is within the safe range. The higher the cumulative probability, the lower the overall risk of the system. The coordination factor is positively correlated with the cumulative probability and is set to a value range of 0 to 1. It is used to balance the contribution ratio of local risk and global risk in the overall system. The weighted expected value of the risk system is obtained by weighted fusion. This means that the weighted expected value of local risk and the weighted expected value of global risk are multiplied by the corresponding coordination factor, and then the product results are added together. The sum is the overall weighted expected value of the risk system. This value is a quantitative representation of the overall risk of the system.

[0116] For example, after normalization, the local risk value is 1, the corresponding weighting coefficient is 0.35, and the weighted expected value of the local risk is 0.35; after normalization, the global risk value is 0, the corresponding weighting coefficient is 0.72, and the weighted expected value of the global risk is 0; the cumulative probability of the coupled state is 95%, based on which the proportion of local risk in the coordination factor is determined to be 0.4, the proportion of global risk is 0.6, and the weighted expected value of the risk system after weighted fusion is 0.14.

[0117] Risk quantification is achieved by calculating the weighted expected value of a single type of risk, and the scientific integration of multi-dimensional risks is completed through coordination factors. The resulting quantified risk value can be used as an optimization target, solving the problem that a single risk cannot represent the overall risk of the system. The weighted expected value of the risk system is the optimization target for subsequent processing.

[0118] The weighted expected value of the risk system only achieves the optimization objective of system risk, without constraining the operational boundaries and coordination rules of scheduling decisions. If the strategy is solved solely with the expected risk value as the objective, decisions may exceed the uncertainty boundary or violate the temporal coordination constraints of coupled states, resulting in the scheduling strategy being unable to be executed. By setting constraint penalty terms in conjunction with the uncertainty boundary and coordination constraints corresponding to the feasible domain of the decision, the uncertainty boundary is treated as a static operational range constraint, and the temporal coordination constraints between coupled states are treated as dynamic correlation constraints. A corresponding penalty function is set for each type of constraint. When the scheduling decision exceeds the uncertainty boundary or violates the coordination constraints, the penalty function outputs a positive penalty value. The higher the degree of decision overshooting or violation, the larger the penalty value. When there is no violation, the penalty value is zero.

[0119] The objective function is constructed based on the weighted expected value of the risk system. This means that the weighted expected value of the risk system is used as the basic optimization term of the objective function, and it is added to the constraint penalty term to form the objective function. The optimization direction of the objective function is to minimize the function value, that is, to simultaneously achieve risk minimization and zero constraint violation. For example, the uncertainty boundary is SOC≤0.85, P≤0.9MW, and the cooperative constraint is that the ratio of the change in P to SOC is between 0.8 and 1.2. The unit penalty value for exceeding the boundary is set to 0.2, and the unit penalty value for violating the cooperative constraint is set to 0.1. Based on the weighted expected value of the risk system of 0.14, the constraint penalty term is added to form the complete objective function. The function value is the sum of the weighted expected value of the risk system and the constraint penalty value.

[0120] This allows the objective function to simultaneously perform both risk optimization and constraint control, ensuring that the subsequent scheduling strategy satisfies both the risk optimization requirement and the system's safe operation rules. The objective function is directly applied to the solution process of the deep deterministic strategy algorithm, and its rationality directly determines the optimization effect and executability of the scheduling strategy.

[0121] Specifically, such as Figure 2 As shown, the initial scheduling policy output includes:

[0122] By combining indicator priority, fluctuation coefficient of coupled state and collaborative constraints, a dual network for deep deterministic strategy algorithm is initialized;

[0123] The samples generated by the interaction between the deep deterministic strategy algorithm and the objective function are obtained, and the samples are stored in a hierarchical manner based on the probability density and the difference in risk level in the risk system. Samples are then sampled from the experience replay pool.

[0124] The dual network is constrained and calibrated based on the uncertainty boundary and cumulative probability. The objective function is then iteratively solved under the constraint calibration. The iteration stops when the solution error of the objective function, the triggering frequency of the constraint penalty term, and the termination condition of the global risk value are all satisfied. The solution result is then output as the initial scheduling policy.

[0125] As a scheduling algorithm based on deep reinforcement learning, the Deep Deterministic Policy Algorithm (DDE) can lead to problems if its policy network and value network are randomly initialized. This can result in the initial parameters of the network being completely disconnected from the risk management logic, parameter fluctuation characteristics, and temporal coordination rules of the energy storage system. Consequently, the algorithm may experience slow iterative convergence, deviation of the initial training direction from the safety constraints, and easy failure during training. The dual networks of the DDE refer to the core policy network and value network. The policy network is used to output the corresponding scheduling action based on the system state, while the value network is used to evaluate the optimization value of the current state and action. Both networks are fully connected deep neural networks.

[0126] The prioritization of local and global risk indicators is transformed into weighted coefficients of dual network input features. High-priority global risk features are assigned higher input weights, while low-priority local risk features are assigned regular input weights, ensuring that the network prioritizes high-level risk management from the initial stage. Furthermore, the volatility coefficient is used as a scaling factor for the initial network weights. The larger the volatility coefficient, the larger the initial weight of the network neuron, thus matching the initial network parameters with the degree of parameter volatility.

[0127] The cooperative constraints between coupled states are transformed into initial boundary constraints of the network output layer, which limit the scheduling actions of the initial output of the network to conform to the parameter cooperative change rules and avoid illegal actions in the initial output. The network initialization also includes setting parameters such as the number of network layers, number of neurons, activation function and basic learning rate, all of which adopt the conventional configuration of deep reinforcement learning in the field of energy storage system scheduling.

[0128] For example, the coupled state is {P, SOC}, consisting of the charging / discharging power P and the state of charge SOC. The global risk is the priority of the first-level indicator, and the local risk is the priority of the second-level indicator, with corresponding input feature weighting coefficients of 1.2 and 1.0, respectively. The fluctuation coefficients of the coupled state are 0.44 and 0.56, respectively, which are used as scaling factors for the initial weights of the policy network and the value network. The collaborative constraint is that the ratio of the change amplitude of P to SOC is between 0.8 and 1.2, and this ratio range is set as the initial constraint of the network output layer. Both the policy network and the value network are set as 3-layer fully connected networks with 64, 32, and 16 neurons, respectively. The activation function is the ReLU function, and the basic learning rate is set to 0.001, completing the initialization of the two networks.

[0129] This avoids the training direction deviation problem caused by random initialization and improves the initial rationality and convergence efficiency of algorithm iteration. The initialized dual network is the carrier of subsequent processing, and the rationality of the initial state of the network determines the quality of sample generation and the effectiveness of subsequent iterative solutions.

[0130] Deep deterministic policy algorithms rely on an experience replay pool to reuse samples. If all samples generated by the interaction between the algorithm and the objective function are randomly mixed and stored, a problem arises where high-risk, low-probability-density critical samples are overwhelmed by a large number of regular samples, preventing the algorithm from fully learning the scheduling logic of high-risk scenarios. Obtaining samples generated by the interaction between the deep deterministic policy algorithm and the objective function refers to the process in which the algorithm takes the current coupling state as input, outputs scheduling actions through a dual network, substitutes the actions into the objective function to obtain a reward value, and then transfers to the next coupling state. The resulting quadruple of state, action, reward, and next state constitutes a single training sample, which is the data updated by the algorithm's network.

[0131] Based on probability density and risk level differences in the risk system, samples are stored in a tiered manner. First, samples are divided into high-density and low-density samples according to the probability density of the coupled state. Samples with a probability density higher than a set threshold are considered high-density normal samples, while those with a probability density lower than the set threshold are considered low-density abnormal samples. Then, samples are divided into local risk samples and global risk samples according to the risk level. Different storage areas of the experience replay pool are divided according to the sample type, and a higher storage ratio is allocated to global risk low-density samples to ensure the storage capacity of key samples. Priority sampling is performed from the experience replay pool according to the tiered storage ratio to increase the sampling probability of global risk low-density samples and avoid training bias caused by insufficient sampling of key samples.

[0132] There are two methods for obtaining the critical value. The first is to calculate the median or upper quartile of the density value based on the probability density distribution of the actual sampled data of the coupled state. The second is to refer to the general critical value range in the field of random quantization of energy storage systems, usually selecting the 20%-30% quartile of the probability density distribution to effectively distinguish between normal operation samples and abnormal operating condition samples. In this embodiment, the critical value is determined by combining statistical methods with engineering experience. First, the probability density distribution of the coupled state {P,SOC} is statistically analyzed, and the density value range is 0.01 to 0.15. The upper quartile is calculated to be 0.1.

[0133] For example, a probability density threshold of 0.1 is set, where samples above this value are considered high-density samples and samples below this value are considered low-density samples. The total storage capacity of the experience replay pool is set to 10,000 samples, with the global risk low-density sample region accounting for 30%, the global risk high-density sample region accounting for 25%, the local risk high-density sample region accounting for 20%, and the local risk low-density sample region accounting for 25%. During sample sampling, the global risk low-density sample sampling ratio is set to 35%, and the other two types of sample sampling ratios combined account for 65%. After completing the stratified sampling, the samples are input into the dual network for parameter updates.

[0134] By implementing differentiated management of samples through hierarchical storage and hierarchical sampling, the training weights of high-risk and low-probability key samples are strengthened, thereby improving the algorithm's learning ability in extreme risk scenarios and ensuring the comprehensiveness of risk control of the scheduling strategy. The sample set is the direct data input for subsequent dual-network constraint calibration and objective function iterative solution. The balance and representativeness of the samples directly affect the accuracy of network calibration and the reliability of iterative solution.

[0135] During the iteration process, the output of the dual network may exceed the uncertainty boundary and violate the cooperative constraints, resulting in the objective function solution not meeting the requirements for safe operation. At the same time, unconstrained iteration will lead to infinite loops. Constraint calibration of the dual network based on the uncertainty boundary and cumulative probability refers to using the previously determined uncertainty boundary as a hard constraint on the dual network output, and performing boundary pruning on the scheduling actions of the network output to ensure that the output result is always within the uncertainty boundary range. At the same time, the cumulative probability is used as a calibration factor for the reward function. When the cumulative probability does not reach the preset safety value, the reward value is reduced to strengthen the network's learning of safety constraints.

[0136] The iterative solution of the objective function under constraint calibration refers to using a dual-network system after constraint calibration as a foundation. The policy network continuously outputs scheduling actions, and the value network continuously evaluates the value of the actions. The network parameters are continuously updated through the backpropagation algorithm, so that the objective function value gradually converges to the optimal value. The termination condition is that the solution error of the objective function, the triggering frequency of the constraint penalty term, and the global risk value all meet the set thresholds. The solution error is the difference between the objective function values ​​of two adjacent iterations. The triggering frequency of the constraint penalty term is the proportion of times the cooperative constraints or uncertainty boundaries are violated during the iteration. The global risk value is the overall risk value of the system after iteration. All three thresholds are set according to the scheduling accuracy and safety management standards of the energy storage system. When all three conditions are met, the iteration stops and the scheduling action output by the current network is used as the initial scheduling strategy.

[0137] For example, the uncertainty boundary is SOC≤0.85, P≤0.9MW, and the cumulative probability is preset to a safety value of 95%, which is used to calibrate the dual network output and reward function. The objective function solution error threshold is set to 0.001, the constraint penalty term trigger frequency threshold is 5%, and the global risk value threshold is 0.05. After 120 iterations, the objective function solution error is 0.0008, the constraint penalty term trigger frequency is 3%, and the global risk value is 0.04. All three values ​​meet the termination condition, so the iteration stops, and the current network output charging and discharging power and state of charge control combination are used as the initial scheduling strategy output.

[0138] By ensuring compliance of the iterative process through constraint calibration and guaranteeing the accuracy and security of the solution results through multi-dimensional termination conditions, the problem of strategy failure caused by unconstrained iteration is solved, and the optimization objective and security constraints are unified. The initial scheduling strategy serves as the basic scheme for the subsequent actual scheduling execution of the system, and its compliance, risk level and optimization degree directly determine the overall effect of energy scheduling of the energy storage system.

[0139] The pre-execution module is used to receive and pre-execute the execution trajectory of the initial scheduling strategy in different time windows, and determine whether the execution trajectory is within the safety boundary. If so, the initial scheduling strategy is output as a scheduling instruction. If not, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation until a scheduling instruction is output.

[0140] Furthermore, the safety boundary is characterized as an operational safety constraint interval composed of the probability distribution of coupled states and local and global risk values. The operational safety constraint interval is based on the uncertainty boundary and uses the temporal correlation between coupled states as a collaborative constraint, which is used to compare the interval affiliation with the execution trajectory.

[0141] If the execution trajectory is within the safety boundary, the initial scheduling policy will be output as a scheduling instruction.

[0142] If the execution trajectory is not within the safety boundary, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation.

[0143] The execution of the initial scheduling strategy relies on clear safety judgment criteria. Relying solely on the uncertainty boundary cannot simultaneously address the requirements of risk control and temporal coordination. However, the probability distribution of coupled states, local risk values, and global risk values ​​can quantify the random characteristics and risk levels of system operation. The temporal correlation between coupled states can define the dynamic coordination rules of parameters. The safety boundary is characterized as an operational safety constraint interval composed of the probability distribution of coupled states and local and global risk values. This means that the safety boundary is not a single fixed numerical boundary, but a multi-dimensional constraint interval that integrates random characteristics and risk levels. The probability distribution of coupled states is used to limit the range adaptability of the interval, ensuring that the interval conforms to the random fluctuation characteristics of the actual system operation.

[0144] Local and global risk values ​​are used to define the safety level of an interval, ensuring that the operational state within the interval meets risk control requirements. Together, they constitute the operational safety constraint interval. The determined uncertainty boundary serves as the basic range of the safety boundary. The interval range of the safety boundary always remains within the uncertainty boundary and does not exceed the system operation limits defined by the uncertainty boundary, ensuring the rationality and safety of the safety boundary. The temporal correlation rules of the coupled state are transformed into dynamic constraints of the safety boundary, limiting the coordinated change patterns of each component of the coupled state within the operational safety constraint interval, avoiding operational states where parameter changes are disconnected or temporal correlations are violated. The purpose of the safety boundary is to compare the interval affiliation with the execution trajectory of the initial scheduling strategy, and to determine whether the execution of the initial scheduling strategy meets safety requirements through comparison.

[0145] For example, the coupled state is {P, SOC}, which consists of the charging / discharging power P and the state of charge (SOC). The probability distribution of the coupled state is {P ∈ [0.2MW, 0.9MW], SOC ∈ [0.3, 0.85]}, with the local risk value limited to ≤0.05 and the global risk value limited to ≤0.03. Based on the uncertainty boundary {P ≤ 0.9MW, SOC ≤ 0.85}, and with the temporal collaborative constraint "the ratio of the change in P to SOC is between 0.8 and 1.2" as the dynamic constraint, an operational safety constraint interval is constructed. That is, the safety boundary is {P ∈ [0.2MW, 0.9MW], SOC ∈ [0.3, 0.85], the range of the ratio of the change in P to SOC is [0.8, 1.2], the local risk value is ≤0.05, and the global risk value is ≤0.03}. This interval is used for the subsequent comparison of the execution trajectory.

[0146] This enables a deep fit between security constraints and the system's random characteristics, risk level, and temporal coordination rules, solving the problem that a single boundary cannot accommodate multi-dimensional security requirements. The security boundary serves as the basis for subsequent processing decisions, and its rationality determines the accuracy of security decisions and the security of subsequent scheduling.

[0147] The initial scheduling strategy is only a theoretical solution for the algorithm, and its actual execution trajectory may deviate from the safety requirements. The execution trajectory refers to the continuous data sequence formed by the changes of each component of the coupled state over time during the actual execution of the initial scheduling strategy. The value of the coupled state at each time node constitutes a data point on the trajectory. Interval attribution comparison refers to judging each data point on the execution trajectory at each time node to see if it simultaneously satisfies all the constraints of the safety boundary, that is, judging whether the value of each component of the coupled state is within the interval of the safety boundary, whether it satisfies the temporal coordination constraint, and whether the corresponding local risk value and global risk value meet the safety boundary limit.

[0148] If all trajectory data points at all time points satisfy the safety boundary constraints, the execution trajectory is determined to be within the safety boundary. In this case, there is no need to adjust the initial scheduling strategy; it can be directly output as the final scheduling command to the execution module of the energy storage system for actual energy scheduling. For example, the execution trajectory of the charging / discharging power P and state of charge (SOC) output by the initial scheduling strategy changes over time in the sequence {P=0.5MW, SOC=0.6}, {P=0.6MW, SOC=0.65}, {P=0.7}. MW, SOC=0.72}, interval assignment comparison is performed node by node. The P value of each node is in [0.2MW, 0.9MW], and the SOC value is in [0.3, 0.85]. The ratio of the change magnitude of P to SOC is 1, 1.08 and 1.14 respectively, all in [0.8, 1.2]. The corresponding local risk value is 0.04 and the global risk value is 0.02, both of which meet the safety boundary limit. It is determined that the execution trajectory is within the safety boundary, and the initial scheduling strategy is output as the scheduling instruction for execution.

[0149] This allows for the feasibility verification of the initial scheduling strategy, ensuring that the output scheduling instructions meet safety requirements and avoiding system operation risks caused by unauthorized scheduling. The initial scheduling strategy is output as a scheduling instruction and directly applied to the actual operation and scheduling of the energy storage system. It is the core execution link of the entire energy scheduling process, and the rationality of its output instructions directly determines the safety and stability of the system operation.

[0150] If the execution trajectory is not within the safety boundary, it indicates that the constraints corresponding to the initial scheduling strategy do not match the actual operating state of the system, or that there is a deviation in the initial constraint settings. The degree of trajectory deviation refers to the distance between the execution trajectory data points and the safety boundary, including the numerical deviation of the coupled state component values ​​exceeding the safety boundary range, the deviation of violating the timing coordination constraints, and the deviation of the risk value exceeding the safety boundary limit. The degree of deviation is quantified by calculating the ratio of the absolute value of the deviation value to the safety boundary range.

[0151] Based on the degree of trajectory deviation, the range of the safety boundary, the parameters of temporal coordination constraints, and the limiting standards for local and global risk values ​​are adjusted. The greater the deviation, the greater the correction magnitude. The correction direction is to reduce the deviation and make the constraints more in line with the actual operating state of the system. The corrected constraints include the range of safety boundary limits, the ratio range of temporal coordination constraints, and the threshold of risk value limits. After correction, it is ensured that the safety boundary is still within the uncertainty boundary and does not exceed the system's operating limits.

[0152] For example, if the execution trajectory of the initial scheduling strategy deviates, and the trajectory data at a certain time node is {P=0.95MW, SOC=0.87}, then at this node, P exceeds the upper limit of the safety boundary by 0.9MW and SOC exceeds the upper limit of the safety boundary by 0.85, with deviations of 5.6% and 2.4% respectively. At the same time, the ratio of the change in P to SOC is 1.3, violating the timing coordination constraint (0.8-1.2), with a deviation of 8.3%. Based on this deviation, the constraints are corrected in reverse, adjusting the safety boundary to {P∈[0.2MW,0.92MW], SOC∈[0.3,0.86]}, and the timing coordination constraint to "the ratio of the change in P to SOC is between 0.75 and 1.25". The local risk value limit is adjusted to ≤0.06, and the global risk value limit is adjusted to ≤0.04. The corrected constraints better match the fluctuation characteristics of the actual system operation and can prevent subsequent execution trajectories from deviating again.

[0153] This solves the problem of mismatch between initial constraints and actual system operation, and improves the adaptability of constraints. The corrected constraints will be fed back to the preceding steps to adjust safety boundaries and optimize the initial scheduling strategy, providing a basis for the iterative optimization of subsequent scheduling strategies, ensuring that the subsequent execution trajectory meets safety requirements, and guaranteeing the long-term safe and stable operation of the system.

[0154] Specifically, the reverse correction of uncertainty boundaries includes:

[0155] The system receives the initial scheduling policy, divides time windows based on the temporal correlation of the coupling state, performs temporal deduction on the changes in the coupling state of the initial scheduling policy in different time windows, and obtains the corresponding execution trajectory.

[0156] If the execution trajectory is not located within the safety boundary, the degree of trajectory deviation is determined based on the spatial deviation of the execution trajectory relative to the safety boundary, the risk deviation value, and the duration of the temporal deviation.

[0157] The quantiles corresponding to the uncertainty boundary are corrected in reverse according to the degree of trajectory deviation, wherein the correction direction of the uncertainty boundary is opposite to the direction of trajectory deviation, and the correction amount is positively correlated with the degree of trajectory deviation.

[0158] After the correction is completed, the timing of the initial scheduling strategy is re-performed and the execution trajectory is compared until the execution trajectory is within the safety boundary, and then the scheduling instruction is output.

[0159] The initial scheduling strategy is a static parameter combination, which cannot directly reflect the dynamic changes of the strategy during actual operation. The coupled state has inherent temporal correlation characteristics, and the parameter changes will show a continuous linkage pattern over time. If the static strategy is directly used for security verification, the real execution state cannot be restored. Receiving the initial scheduling strategy refers to obtaining the initial scheduling strategy output by the preceding deep deterministic strategy algorithm. This strategy contains the target running parameters of each component of the coupled state and is the basic input object for temporal deduction.

[0160] Based on the linkage change cycle between the components of the coupled state and the regular scheduling cycle of the energy storage system, the complete scheduling period is divided into several continuous time windows of equal length. The length of the time window is consistent with the response cycle of the time-series correlation of the coupled state, ensuring that the simulation results conform to the actual operation law. According to the inherent time-series coordination constraints of the coupled state, the parameter change values ​​of each component of the coupled state are derived window by window. The simulation results of each time window are arranged in chronological order to form a complete continuous execution trajectory. This execution trajectory is an ordered data sequence of the parameters of the coupled state changing with time.

[0161] For example, the initial scheduling strategy received is a coupled state control parameter consisting of charging / discharging power P and state of charge (SOC). The temporal correlation of the coupled state is that the ratio of the change amplitude of charging / discharging power to that of SOC is between 0.8 and 1.2. Based on the scheduling characteristics of the energy storage system, a 5-minute interval is divided into a single time window. The initial scheduling strategy is continuously extrapolated for 3 time windows, resulting in the execution trajectory data sequence as follows: first window {P=0.5MW, SOC=0.6}, second window {P=0.7MW, SOC=0.7}, and third window {P=0.95MW, SOC=0.87}.

[0162] This fully restores the parameter change process of the strategy in actual operation, eliminating the problem of disconnect between static verification and actual operation; the execution trajectory obtained by time-series deduction is the data basis for subsequent processing, and the authenticity and completeness of the trajectory determines the accuracy of the deviation judgment result.

[0163] Simply determining whether the execution trajectory is within the safety boundary cannot provide a quantitative basis for subsequent uncertainty boundary correction. Vague deviation judgments can lead to a lack of specificity in the correction process, resulting in over-correction or under-correction. The derived execution trajectory is compared with the safety boundary window by window. If the coupling state parameters, risk values, or temporal coordination relationships in any window do not meet the safety boundary constraints, it is determined as a trajectory deviation. The degree of trajectory deviation is determined based on the spatial deviation of the execution trajectory relative to the safety boundary, the risk deviation value, and the duration of the temporal deviation.

[0164] Spatial deviation refers to the difference between the actual value of each component of the coupled state in the execution trajectory and the corresponding interval limit of the safety boundary, which is used to characterize the extent of the parameter's out-of-bounds behavior in the spatial dimension; risk deviation refers to the difference between the local risk value, global risk value and the safety boundary risk limit threshold corresponding to the execution trajectory, which is used to characterize the extent of exceeding the limit in the risk dimension; temporal deviation duration refers to the total duration of the time window during which the execution trajectory continuously violates the temporal coordination constraints of the coupled state, which is used to characterize the duration of the violation in the temporal dimension; the degree of trajectory deviation is the weighted comprehensive value of the above three types of deviation indicators, with the weights allocated according to the priority of safety control, and the weights of spatial deviation and risk deviation are higher than those of temporal deviation duration.

[0165] For example, the safety boundary is defined as charging / discharging power P≤0.9MW, state of charge (SOC)≤0.85, global risk value≤0.03, and timing coordination ratio of 0.8-1.2. The third window of the execution trajectory {P=0.95MW, SOC=0.87} exceeds the safety boundary. The calculated spatial deviation is P exceeding 0.05MW, SOC exceeding 0.02, and risk deviation is global risk value exceeding 0.02. The timing deviation duration is a single time window of 5 minutes. After weighted calculation, the trajectory deviation degree is 0.12.

[0166] This transforms the ambiguous deviation state into a calculable degree of deviation, avoiding correction bias caused by subjective judgment and ensuring the accuracy and rationality of subsequent correction processes. The determined degree of trajectory deviation serves as the basis for subsequent reverse correction, determining the rationality of the correction magnitude for uncertainty boundaries.

[0167] The uncertainty boundary is determined by the quantiles corresponding to the preset confidence level. The fundamental reason for the deviation of the execution trajectory is that the original quantile setting does not match the actual operating characteristics of the system. If the quantiles are not corrected accordingly, adjusting only the boundary values ​​will destroy the statistical rationality of the uncertainty boundary. The quantiles corresponding to the uncertainty boundary are the quantiles determined based on the cumulative probability and the preset confidence level. These quantiles are the core statistical basis of the uncertainty boundary. Changes in the quantile values ​​change the range of the uncertainty boundary.

[0168] The reverse correction of quantiles based on the degree of trajectory deviation means that when the trajectory deviation direction exceeds the upper limit of the safety boundary, the quantiles are corrected in a downward direction; when the trajectory deviation direction is below the lower limit of the safety boundary, the quantiles are corrected in an upward direction. That is, the correction direction is completely opposite to the trajectory deviation direction. The correction magnitude of the quantiles is positively correlated with the degree of trajectory deviation. The greater the trajectory deviation, the greater the adjustment magnitude of the quantiles; the smaller the trajectory deviation, the smaller the adjustment magnitude of the quantiles. During the correction process, it is necessary to ensure that the uncertainty boundary is always within the system's operating limit range and does not exceed the safety bottom line.

[0169] For example, the original uncertainty boundary corresponds to the quantile at the 95% confidence level. The trajectory deviates upwards beyond the safety boundary. Therefore, the quantile is corrected in the downward direction. The trajectory deviation is 0.12, which corresponds to correcting the quantile from 95% to 93%. The corrected uncertainty boundary is that the charge / discharge power P ≤ 0.88MW and the state of charge (SOC) ≤ 0.83.

[0170] This enables adaptive adjustment of the uncertainty boundary, ensuring both the statistical rationality of the boundary and adapting it to the fluctuation characteristics of the actual system operation, thus avoiding control failure caused by adjusting the boundary without a basis. The corrected quantile will regenerate a new uncertainty boundary, which provides an updated constraint basis for subsequent re-time series extrapolation and execution trajectory comparison.

[0171] A single uncertainty boundary correction cannot guarantee that the execution trajectory will meet the safety boundary requirements. If there is still a deviation after correction, the correction process needs to be iterated until the execution trajectory falls completely within the safety boundary. After the correction is completed, the initial scheduling strategy is re-performed and the execution trajectory is compared. That is, the safety boundary is updated with the corrected uncertainty boundary, the time window division and time sequence simulation are repeated, the execution trajectory is generated again and compared with the updated safety boundary. If the execution trajectory is still not within the safety boundary, the quantile reverse correction operation is repeated to form a closed-loop iterative process of correction, simulation and comparison.

[0172] When the execution trajectory is verified to be completely within the safety boundary, the iteration loop terminates, and the initial scheduling strategy adapted to the corrected boundary is output as the final scheduling instruction. For example, after quantile correction, the uncertainty boundary is updated to P≤0.88MW and SOC≤0.83. The execution trajectories obtained by re-temporal derivation are {P=0.5MW, SOC=0.6}, {P=0.7MW, SOC=0.7}, and {P=0.88MW, SOC=0.83}. After window-by-window comparison, all trajectory data meet the safety boundary constraints, and it is determined that the execution trajectory is within the safety boundary. The iteration terminates and the final scheduling instruction is output.

[0173] This continuously optimizes the adaptability of uncertainty boundaries and execution trajectories, ensuring that the output scheduling instructions fully comply with safe operation requirements and mitigating the risk of non-compliance in scheduling strategy execution from a process perspective. The scheduling instructions are directly applied to the actual energy scheduling execution of the energy storage system and are the final output of the entire energy scheduling system. The compliance of the instructions determines the safety and stability of the energy storage system operation.

[0174] The feedback module is used to obtain the actual response trajectory of the scheduling instruction and perform time-series deviation analysis with the execution trajectory. Based on the statistical characteristics of the time-series deviation, the objective function and safety boundary are dynamically adjusted.

[0175] Specifically, dynamically adjusting the objective function includes:

[0176] Obtain the actual response trajectory of the scheduling instruction, perform timing deviation analysis between the actual response trajectory and the execution trajectory obtained from the pre-simulation, and extract the statistical characteristics of the timing deviation, including amplitude, frequency and cumulative amount.

[0177] Based on the statistical characteristics of time series deviations, the weighting coefficients, coordination factors, and constraint penalty terms in the objective function are dynamically adjusted.

[0178] The execution trajectory of the initial scheduling strategy under the constraints of the adjusted objective function is simulated, and the statistical characteristics of the timing deviation after adjustment are verified until the verification is passed to determine the parameter configuration of the objective function.

[0179] When the verified dispatch instructions are executed in the actual energy storage system, the actual running trajectory will deviate from the execution trajectory predicted in the previous theoretical simulation due to factors such as system hardware response delay, fluctuations in operating conditions, and external environmental interference. This deviation will directly cause the optimization logic of the objective function to become disconnected from the actual system operating characteristics. Without quantitative analysis and feature extraction, subsequent adjustments to the objective function will lack objective basis and will not be able to adapt the strategy to the actual system.

[0180] The data acquisition unit of the energy storage system collects the actual operating parameters of each component of the coupled state during the execution of the scheduling command in real time, forming a continuous time-series data sequence in chronological order. This sequence is the actual response trajectory that reflects the true response state of the system. Using the same time window as the comparison unit, the numerical difference of the coupled state parameters in the actual response trajectory and the pre-execution trajectory is calculated window by window. This difference is the time-series deviation. The complete statistics of the deviation over the entire time period are achieved through point-by-point comparison.

[0181] The statistical characteristics of timing deviations, including amplitude, frequency, and cumulative amount, are extracted. Amplitude refers to the maximum absolute value of the timing deviation within a single time window, which is used to characterize the extreme degree of a single deviation. Frequency refers to the proportion of time windows in which the timing deviation exceeds the preset deviation threshold to the total number of time windows, which is used to characterize the frequency of deviation occurrence. Cumulative amount refers to the sum of the absolute values ​​of timing deviations within all time windows, which is used to characterize the overall accumulation of deviations over the entire scheduling period. These three types of characteristics together complete the multi-dimensional quantification of timing deviations.

[0182] For example, the coupled state is {P,SOC} consisting of the charging / discharging power P and the state of charge SOC. The pre-execution trajectory and the actual response trajectory corresponding to the scheduling command are both divided into 5-minute time windows. The timing deviation is obtained by comparing the time windows one by one. The amplitude characteristics are that the maximum deviation of charging / discharging power is 0.08MW and the maximum deviation of state of charge is 0.06. The frequency characteristics are that the proportion of windows with deviations exceeding the critical value of 0.03 is 20%. The cumulative characteristic is that the sum of the absolute values ​​of the deviations over the entire time period is 0.32.

[0183] This transforms the ambiguous actual operational deviations into quantifiable statistical characteristics, eliminating the information gap between the theoretical model and the actual system, and providing an objective basis for adjusting the objective function. The extracted statistical characteristics of the time-series deviation amplitude, frequency, and cumulative amount are direct input parameters for subsequent processing, and their quantification accuracy directly determines the pertinence and rationality of the objective function adjustment.

[0184] The weighting coefficients, coordination factors, and constraint penalty terms in the objective function are all set based on the previous theoretical model. The three statistical characteristics of the time series deviation correspond to the adaptation defects of different parameters of the objective function. Excessive amplitude deviation reflects the insufficient adaptation of the weighting coefficients to the fluctuation of the coupled state. Excessive frequency deviation reflects the failure of the coordination factor to balance local and global risks. Excessive cumulative deviation reflects the insufficient constraint of the constraint penalty term on the violation.

[0185] Based on the statistical characteristics of time-series deviations, the weighting coefficients, coordination factors, and constraint penalty terms in the objective function are dynamically adjusted. The weighting coefficients are corrected coefficients used for risk weighting calculations. When the magnitude of the time-series deviation exceeds a set range, the weighting coefficients are adjusted proportionally to the deviation magnitude, strengthening the weight control of high-deviation parameters. The coordination factor is a proportional factor used to balance local and global risks. When the frequency of the time-series deviation exceeds a set range, the allocation ratio of the coordination factor is adjusted to optimize the balance between the two types of risks. The constraint penalty term is a penalty value used to constrain violations. When the cumulative amount of time-series deviation exceeds a set range, the value of the constraint penalty term is increased to strengthen the constraint on deviation behavior. All parameter adjustments are based solely on the actual deviation characteristics and are not subject to subjective or arbitrary adjustments.

[0186] For example, the original objective function weighting coefficients were 0.35 for local risk and 0.72 for global risk. Because the magnitude deviation exceeded the threshold, the weighting coefficients were adjusted to 0.38 for local risk and 0.75 for global risk. The original coordination factor was 0.4 for local risk and 0.6 for global risk. Because the frequency deviation reached 20%, the coordination factor was adjusted to 0.35 for local risk and 0.65 for global risk. The original constraint penalty term unit value was 0.1. Because the cumulative deviation reached 0.32, the constraint penalty term unit value was increased to 0.15.

[0187] This enables dynamic adaptation between the objective function and the actual system operating characteristics, improving the authenticity and reliability of the objective function optimization results. The adjusted weighting coefficients, coordination factors, and constraint penalty terms constitute the updated parameter configuration of the objective function, which is the basis for subsequent processing.

[0188] Adjusting the objective function parameters based on timing deviations in a single instance cannot guarantee that the adjusted parameter configuration will keep the timing deviations within a reasonable range. If the adjusted parameters are used directly, the deviation may still exceed the limit. Pre-simulating the execution trajectory of the initial scheduling strategy under the constraints of the adjusted objective function means using the updated objective function parameter configuration, repeating the timing deduction process, and theoretically pre-simulating the initial scheduling strategy again to obtain a new execution trajectory adapted to the adjusted objective function.

[0189] The statistical characteristics of the timing deviation after verification and adjustment refer to performing timing deviation analysis again on the execution trajectory obtained from the new pre-simulation and the actual response trajectory, re-extracting three types of statistical characteristics: amplitude, frequency, and cumulative amount, and determining whether all three characteristics are within the preset reasonable threshold range. This threshold range is set according to the energy storage system scheduling accuracy standard. If the statistical characteristics meet the threshold requirements, the verification is deemed successful, and the currently adjusted weighting coefficient, coordination factor, and constraint penalty term are used as the parameters of the objective function. If the verification fails, the parameter adjustment operation of the previous step is repeated until the statistical characteristics meet the requirements.

[0190] For example, the preset timing deviation verification thresholds are amplitude ≤ 0.03, frequency ≤ 10%, and cumulative amount ≤ 0.1. The execution trajectory obtained by pre-simulating the adjusted objective function is used. After deviation analysis, the amplitude is 0.02, the frequency is 8%, and the cumulative amount is 0.08. All three statistical characteristics meet the threshold requirements, the verification is passed, and the adjusted parameters are determined to be the final configuration of the objective function.

[0191] This ensures the optimality and stability of the objective function parameter configuration, eliminates the mismatch between theory and practice, and ensures that the objective function can continuously output a scheduling strategy that meets the actual operation requirements. The parameter configuration of the objective function is used in the subsequent continuous energy scheduling process of the energy storage system, serving as the calculation basis for the iterative optimization of the scheduling strategy, and directly determining the accuracy and stability of the system's long-term scheduling.

[0192] Specifically, dynamically adjusting the safety boundary includes:

[0193] Based on the statistical characteristics of the time-series deviation, and combined with the probability distribution, local risk value, and global risk value of the coupling state corresponding to the safety boundary, the interval threshold and cooperative constraints of the safety boundary are dynamically adjusted.

[0194] The adjusted safety boundary and uncertainty boundary are subjected to constraint consistency verification, the execution trajectory of the initial scheduling strategy is rehearsed, and the interval attribution comparison results are verified until the verification is passed to determine the safety boundary.

[0195] After actual scheduling and execution verification, the statistical characteristics of timing deviation can intuitively reflect the mismatch between the theoretically set safety boundary and the actual operating conditions of the energy storage system. The original safety boundary is only constructed based on the theoretical probability distribution of the coupled state, local risk value and global risk value, without incorporating the actual operating deviation characteristics. This can easily lead to overly strict safety boundary constraints, which restricts the scheduling strategy, or overly loose constraints, which lead to the failure of safety control.

[0196] The statistical characteristics of time-series deviations are three types of features: amplitude, frequency, and cumulative amount. Among them, the amplitude feature represents the extreme value of a single deviation between the actual response and the simulated trajectory, the frequency feature represents the frequency of deviation occurrence, and the cumulative amount represents the overall scale of deviation over the entire time period. The interval threshold of the safety boundary refers to the upper and lower limits of the range of values ​​of the coupled state, which is the static spatial constraint of the safety boundary. The collaborative constraint refers to the parameter linkage change rules formed based on the temporal correlation of the coupled state, which is the dynamic rule constraint of the safety boundary.

[0197] The dynamic adjustment process requires correlating and matching the statistical characteristics of time-series deviations with the probability distribution of coupled states, and the local risk values ​​with the global risk values. When the magnitude of the time-series deviation exceeds a reasonable range, the upper and lower limits of the interval threshold are adjusted accordingly. The larger the deviation, the larger the adjustment range of the interval threshold. When the frequency of time-series deviations is too high, the linkage range of the parameters of the collaborative constraints is relaxed or narrowed accordingly. When the cumulative amount of time-series deviation exceeds the standard, the interval threshold and collaborative constraints are simultaneously fine-tuned in combination with the quantile characteristics of the probability distribution and the control level of the risk value. During the adjustment process, the ability of the safety boundary to control the system operation risk is always maintained, and the limit standards of local risk and global risk are not reduced.

[0198] For example, the timing deviation statistical characteristics are a charge / discharge power amplitude deviation of 0.08MW, a state of charge amplitude deviation of 0.06, a deviation frequency ratio of 20%, and a cumulative deviation of 0.32; the coupled state probability distribution corresponds to the 93% confidence level quantile, with local risk values ​​limited to no more than 0.05 and global risk values ​​limited to no more than 0.03; the original safety boundary interval thresholds are charge / discharge power P∈[0.2MW, 0.88MW], state of charge SOC∈[0.3, 0.83], and the cooperative constraint is that the ratio of P to SOC change amplitude ranges from [0.8, 1.2]; combining the deviation characteristics with risk and probability distribution parameters, the interval thresholds are adjusted to P∈[0.2MW, 0.90MW] and SOC∈[0.3, 0.85], and the cooperative constraint is adjusted to the ratio of P to SOC change amplitude ranges from [0.75, 1.25]. After adjustment, the risk control requirements of local risk value ≤ 0.05 and global risk value ≤ 0.03 are still maintained.

[0199] This eliminates the mismatch between the safety boundary and the actual working conditions, and improves the adaptability and rationality of the safety boundary while ensuring risk control capabilities. The adjusted safety boundary interval thresholds and collaborative constraints provide adjustment objects for subsequent processing and are the basic input conditions for the subsequent closed-loop verification process.

[0200] The dynamically adjusted security boundary must use the uncertainty boundary as the underlying constraint benchmark. If the adjusted security boundary exceeds the limit of the uncertainty boundary, it will disrupt the underlying security logic of the system operation. At the same time, the adjusted security boundary needs to verify its constraint effectiveness on the execution trajectory through actual pre-running. The constraint consistency verification between the adjusted security boundary and the uncertainty boundary refers to verifying whether the threshold of the adjusted security boundary interval is completely within the interval range of the uncertainty boundary, and at the same time verifying whether there is a conflict between the collaborative constraints of the security boundary and the statistical rules of the uncertainty boundary. The uncertainty boundary is the system operation limit boundary determined based on the probability distribution quantile. The constraint consistency verification is the core link to ensure that the security boundary does not break through the underlying security bottom line.

[0201] Using the same time-series extrapolation method as described above, based on the adjusted safety boundary and time window division rules, the initial scheduling strategy is extrapolated throughout the entire time period to obtain the corresponding theoretical execution trajectory. The execution trajectory obtained from the pre-simulation is compared with the adjusted safety boundary on a time-window basis to determine whether all data points of the trajectory are within the safety boundary constraints. If there is a trajectory out of bounds, the safety boundary adjustment operation is repeated. If all trajectory data meet the safety boundary constraint requirements, the verification is considered successful. Through the closed-loop process of iterative adjustment and verification, the safety boundary that meets the requirements of constraint consistency and trajectory attribution is taken as the final determined safety boundary, which is used for the verification and execution of subsequent scheduling strategies.

[0202] For example, if the uncertainty boundary is that the charging / discharging power P≤0.92MW and the state of charge (SOC)≤0.86MW, the adjusted safety boundary interval threshold is P∈[0.2MW, 0.90MW] and SOC∈[0.3, 0.85]. After verification, the safety boundary is within the uncertainty boundary range, and there are no rule conflicts in the cooperative constraints. The execution trajectory obtained after the time-series deduction of the initial scheduling strategy is within the adjusted safety boundary after window-by-window comparison. The interval assignment comparison verification is passed, and the adjusted safety boundary is determined as the final safety boundary.

[0203] Through a dual closed-loop mechanism of constraint consistency verification and trajectory pre-simulation verification, the compliance and practicality of the adjusted safety boundary are ensured. This ensures that the underlying uncertainty boundary constraints are not exceeded, while effectively constraining the scheduling execution trajectory, thus achieving dynamic optimization and stable implementation of the safety boundary. The determined safety boundary will serve as the basis for subsequent scheduling strategy verification, execution trajectory comparison, and constraint condition correction of the energy storage system. At the same time, it will be fed back to the dynamic adjustment of the objective function, forming a constraint closed loop for the entire scheduling system process, providing a standardized constraint foundation for the long-term safe and stable operation of the system.

[0204] Example 2:

[0205] like Figure 3 The diagram shown is a flowchart of an energy dispatching method for an energy storage system provided in this application embodiment. The method includes:

[0206] Acquire state data reflecting the operating status of the energy storage system, quantify the probability distribution of the state data through multinomial chaos expansion, construct a risk system based on the probability distribution, and generate a decision feasible region including uncertainty boundaries and risk weights.

[0207] The objective function is constructed by the weighted expected value of the uncertainty boundary and risk system corresponding to the decision feasible region. The objective function is solved by the deep deterministic strategy algorithm to output the initial scheduling strategy.

[0208] Receive and rehearse the execution trajectory of the initial scheduling strategy within different time windows, and determine whether the execution trajectory is within the safety boundary. If so, output the initial scheduling strategy as a scheduling instruction. If not, reverse the uncertainty boundary according to the degree of trajectory deviation until the scheduling instruction is output.

[0209] The actual response trajectory of the scheduling instruction is obtained, and a timing deviation analysis is performed between the trajectory and the execution trajectory. Based on the statistical characteristics of the timing deviation, the objective function and safety boundary are dynamically adjusted.

[0210] Since the principle of the method in this application embodiment is similar to that of the system described in this application embodiment, the implementation of the method is the same as that of the system, and the repeated parts will not be described again.

Claims

1. An energy dispatching system for an energy storage system, characterized in that, include: The module includes a fusion module, a scheduling module, a pre-rendering module, and a feedback module. The fusion module is used to acquire state data reflecting the operating status of the energy storage system, quantifies the probability distribution of the state data through multinomial chaos expansion, constructs a risk system based on the probability distribution, and generates a decision feasible domain including uncertainty boundaries and risk weights. The scheduling module is used to construct an objective function based on the weighted expected value of the uncertainty boundary and risk system corresponding to the decision feasible region. The objective function is solved by a deep deterministic strategy algorithm to output the initial scheduling strategy. The pre-execution module is used to receive and pre-execute the execution trajectory of the initial scheduling strategy in different time windows, and determine whether the execution trajectory is within the safety boundary. If so, the initial scheduling strategy is output as a scheduling instruction. If not, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation until the scheduling instruction is output. The feedback module is used to obtain the actual response trajectory of the scheduling instruction and perform time-series deviation analysis with the execution trajectory. Based on the statistical characteristics of the time-series deviation, the objective function and safety boundary are dynamically adjusted.

2. The energy dispatching system for an energy storage system as described in claim 1, characterized in that, The probability distribution of the quantized state data includes: Obtain state data reflecting the operating status of the energy storage system, calculate the contribution of the state data to the uncertainty of the energy storage system operation using the mutual information method, filter out state data with a contribution greater than the contribution threshold, and form coupled states through time series correlation. The mean square error between the empirical distribution and the standard random distribution of the coupled state is calculated. The basis functions are matched, and the deceleration rate of the chaotic expansion error is used as the convergence exponent until it is less than the convergence threshold. Then the expansion order is output. The coupled state is expanded using a polynomial chaotic expansion with basis functions and expansion order. The expansion coefficients are solved by minimum angle regression. After removing terms whose absolute values ​​of the expansion coefficients are less than a preset threshold, a sparse chaotic expansion model is obtained. The probability distribution, including probability density and cumulative probability, in the coupled state is reconstructed based on the sparse chaotic expansion model, and the quantile corresponding to the preset confidence level is used as the uncertainty boundary.

3. The energy dispatching system for an energy storage system as described in claim 2, characterized in that, Generating the decision feasible domain includes: Based on the probability distribution including probability density and cumulative probability in the coupled state, a risk system including local risk and global risk is constructed by dividing the scope of influence. The volatility coefficient of the coupled state is determined based on the variance and covariance of the probability distribution. Risk weights are allocated according to the magnitude of the volatility coefficients, and the risk weights are adjusted according to the differences in risk levels in the risk system. Based on uncertainty boundaries and risk weights, and combined with the collaborative constraints formed by the temporal correlation between coupled states, a decision-feasible region is generated.

4. The energy dispatching system for an energy storage system as described in claim 3, characterized in that, The objective function comprises: The local and global risk values ​​are normalized, and the priority of the indicators for local and global risks is determined by combining the magnitude of the volatility coefficient of the coupled state. The weighting coefficients are obtained by calibrating the initial coefficients using the corrected risk weights and combining them with the variance and covariance of the probability distribution. The weighting coefficients are then adjusted according to the priority of the indicators. The normalized local risk value, the normalized global risk value, and the corresponding modified weighting coefficient are multiplied to obtain the weighted expected values ​​of local risk and global risk, respectively. The coordination factor is determined according to the magnitude of the cumulative probability, and the weighted expected value of the risk system is obtained by weighted fusion. Based on the weighted expected value of the risk system, and combined with the constraint penalty terms set by the uncertainty boundary and collaborative constraints corresponding to the feasible decision domain, the objective function is constructed.

5. The energy dispatching system for an energy storage system as described in claim 4, characterized in that, The initial output scheduling strategy includes: By combining indicator priority, fluctuation coefficient of coupled state and collaborative constraints, a dual network for deep deterministic strategy algorithm is initialized; The samples generated by the interaction between the deep deterministic strategy algorithm and the objective function are obtained. The samples are stored in a hierarchical manner based on the probability density and the difference in risk level in the risk system, and samples are sampled from the experience replay pool. The dual network is constrained and calibrated based on the uncertainty boundary and cumulative probability. The objective function is then iteratively solved under the constraint calibration. The iteration stops when the solution error of the objective function, the triggering frequency of the constraint penalty term, and the termination condition of the global risk value are all satisfied. The solution result is then output as the initial scheduling policy.

6. The energy dispatching system for an energy storage system as described in claim 5, characterized in that, The safety boundary is characterized as an operational safety constraint interval composed of the probability distribution of coupled states and local and global risk values. The operational safety constraint interval is based on the uncertainty boundary and uses the temporal correlation between coupled states as a collaborative constraint, which is used to compare the interval affiliation with the execution trajectory. If the execution trajectory is within the safety boundary, the initial scheduling policy will be output as a scheduling instruction. If the execution trajectory is not within the safety boundary, the uncertainty boundary is corrected in reverse according to the degree of trajectory deviation.

7. The energy dispatching system for an energy storage system as described in claim 6, characterized in that, The reverse correction uncertainty boundary includes: The system receives the initial scheduling policy, divides time windows based on the temporal correlation of the coupling state, performs temporal deduction on the changes in the coupling state of the initial scheduling policy in different time windows, and obtains the corresponding execution trajectory. If the execution trajectory is not located within the safety boundary, the degree of trajectory deviation is determined based on the spatial deviation of the execution trajectory relative to the safety boundary, the risk deviation value, and the duration of the temporal deviation. The quantiles corresponding to the uncertainty boundary are corrected in reverse according to the degree of trajectory deviation, wherein the correction direction of the uncertainty boundary is opposite to the direction of trajectory deviation, and the correction amount is positively correlated with the degree of trajectory deviation. After the correction is completed, the timing of the initial scheduling strategy is re-performed and the execution trajectory is compared until the execution trajectory is within the safety boundary, and then the scheduling instruction is output.

8. The energy dispatching system for an energy storage system as described in claim 7, characterized in that, The dynamically adjusted objective function includes: Obtain the actual response trajectory of the scheduling instruction, perform timing deviation analysis between the actual response trajectory and the execution trajectory obtained from the pre-simulation, and extract the statistical characteristics of the timing deviation, including amplitude, frequency and cumulative amount. Based on the statistical characteristics of time series deviations, the weighting coefficients, coordination factors, and constraint penalty terms in the objective function are dynamically adjusted. The execution trajectory of the initial scheduling strategy under the constraints of the adjusted objective function is simulated, and the statistical characteristics of the timing deviation after adjustment are verified until the verification is passed to determine the parameter configuration of the objective function.

9. The energy dispatching system for an energy storage system as described in claim 8, characterized in that, Dynamically adjusting the security boundary includes: Based on the statistical characteristics of the time-series deviation, and combined with the probability distribution, local risk value, and global risk value of the coupling state corresponding to the safety boundary, the interval threshold and cooperative constraints of the safety boundary are dynamically adjusted. The adjusted safety boundary and uncertainty boundary are subjected to constraint consistency verification, the execution trajectory of the initial scheduling strategy is rehearsed, and the interval attribution comparison results are verified until the verification is passed to determine the safety boundary.

10. An energy dispatching method for an energy storage system, implemented based on an energy dispatching system for an energy storage system according to any one of claims 1-9, characterized in that, include: Acquire state data reflecting the operating status of the energy storage system, quantify the probability distribution of the state data through multinomial chaos expansion, construct a risk system based on the probability distribution, and generate a decision feasible region including uncertainty boundaries and risk weights. The objective function is constructed by the weighted expected value of the uncertainty boundary and risk system corresponding to the decision feasible region. The objective function is solved by the deep deterministic strategy algorithm to output the initial scheduling strategy. Receive and rehearse the execution trajectory of the initial scheduling strategy within different time windows, and determine whether the execution trajectory is within the safety boundary. If so, output the initial scheduling strategy as a scheduling instruction. If not, reverse the uncertainty boundary according to the degree of trajectory deviation until the scheduling instruction is output. The actual response trajectory of the scheduling instruction is obtained, and a timing deviation analysis is performed between the trajectory and the execution trajectory. Based on the statistical characteristics of the timing deviation, the objective function and safety boundary are dynamically adjusted.