Knowledge embedding-based hybrid distribution risk assessment method and system for power distribution network scheduling
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
在风险建模方面,现有的约束马尔可夫决策过程框架多基于期望型风险度量进行建模,难以有效刻画电压越限等约束的非对称尾部风险特性,无法实现配电网实时运行过程中经济性与安全性的协同优化
本发明一方面提出一种基于分布鲁棒机会约束的参考调度知识嵌入的残差策略架构,将分布鲁棒优化求解得到的参考调度信息引入强化学习过程,以增强策略学习的可行性引导能力与收敛效率;另一方面提出一种混合分布风险建模方法,用于刻画传统基于期望的风险度量难以表达的尾部风险特性,从而实现对电压越限等安全约束非对称风险的精细建模与约束表达,从而在考虑源荷不确定性的条件下,实现配电网实时运行过程中经济性与安全性的协同优化。
Smart Images

Figure CN122549804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system and artificial intelligence interdisciplinary technology, specifically a hybrid distributed risk assessment distribution network dispatching method and system based on knowledge embedding. Background Technology
[0002] With the advancement of the "dual carbon" target and the accelerated construction of new power systems, the penetration rate of renewable energy in distribution networks is continuously increasing, and the operating environment of distribution networks is exhibiting significant uncertainties. Against this backdrop, real-time operation of distribution networks not only needs to meet economic requirements but also strictly adhere to operational constraints such as voltage, placing higher demands on dispatching decisions.
[0003] Traditional analytical optimization methods suffer from high computational complexity, making them unsuitable for online solution requirements in real-time scheduling scenarios. Reinforcement learning frameworks based on Constrained Markov Decision Processes (CMMs) can introduce safety constraints while optimizing economic objectives, thus enabling policy learning under constraints. However, in complex operational constraints and uncertainties, traditional safety reinforcement learning methods are prone to suboptimal results. To improve state observability, existing methods often employ spatiotemporal graph neural networks for feature extraction; however, these methods are essentially data-driven black-box models, their performance highly dependent on the quality of training data, and their generalization ability is limited. Knowledge embedding-based methods primarily focus on ensuring single-step instantaneous feasibility constraints, failing to fully characterize the temporal coupling constraints and future uncertainties during operation. Regarding risk modeling, existing CMM frameworks are mostly based on expected risk metrics, making it difficult to effectively characterize the asymmetric tail risk characteristics of constraints such as voltage limit exceedances, and thus unable to achieve coordinated optimization of economy and safety during real-time distribution network operation. Summary of the Invention
[0004] To address the shortcomings mentioned in the background section, the present invention aims to provide a method and system for distribution network scheduling based on knowledge embedding for hybrid distributed risk assessment.
[0005] Firstly, the objective of this invention can be achieved through the following technical solution: a hybrid distributed risk assessment and distribution network dispatching method based on knowledge embedding, the method comprising the following steps: The distribution network operation status data is acquired and input into a pre-established intraday rolling time-domain dynamic programming model. By introducing bibliometric chance constraints, a reference dispatch action is output. The intraday rolling time-domain dynamic programming model is used to characterize the evolution of the distribution network operation status and the dispatch decision process under multiple time scales, and a constrained Markov decision process is constructed based on the dynamic programming model. The distribution network operation status data and reference scheduling actions are input into a pre-established residual strategy learning network model, and the adjustment amount of the reference scheduling actions is output. Based on the adjustment amount of the reference scheduling actions, the reference scheduling actions are adjusted and optimized to obtain the adjusted scheduling actions. The cumulative voltage violation cost is obtained, and then input into a pre-established hybrid distribution model. Based on the hybrid distribution model, the conditional value of risk is calculated using closed-loop analytical methods, and the tail risk assessment result is output. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. Based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints, the strategy is solved by co-training the strategy network and the risk assessment network to obtain the optimal scheduling strategy that meets the safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are optimized in a coordinated manner. The tail risk assessment results are used as a hybrid distributed risk constraint.
[0006] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the intraday rolling time-domain dynamic programming model of the distribution network extracts the state transition relationship of a single time segment, describes the operation process of the distribution network as a constrained Markov decision process, and defines state variables, action variables, and a reward function based on the single-step operation cost. By describing the optimal scheduling process of the distribution network at multiple time scales, the scheduling scope is... ,as follows: in, This is a function of total operating cost; This is the cost function for a single-step operation. The power flow equation is as follows; For operational constraints; and In uncertainty Risk operators defined above; These are decision variables, defined within the feasible region determined by the operational constraints. This includes the active power output of distributed generators and the charging and discharging power of energy storage devices; It is a state variable.
[0007] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the state variable comprising active power of the baseline load demand. and reactive power Active power of renewable energy units Remaining electricity of energy storage devices and the price of electricity transmission from substations ,as follows: in, For the state variables of the decision-making agent; The action variables are defined on nodes equipped with distributed generators and energy storage devices, as follows: in, A collection of nodes equipped with energy storage devices; Action variables generated for the decision-making agent; For the action variables corresponding to the output of the distributed generators at each node; Action variables corresponding to the charging and discharging power of energy storage devices at each node; The reward function based on the single-step running cost is as follows: Where, r t For agent reward variables; This represents the active power price of the power grid at time t; This indicates the active power transmitted from the upper-level power grid to the local power grid; This represents the active power quotation of the distributed generator connected to node n; This represents the active power output of the distributed generator connected to node n; This represents the set of nodes connected to distributed generators in the power grid; The cost function takes into account voltage safety costs, as follows: in, For voltage safety costs; To be a function that takes positive values; and These represent the minimum and maximum permissible voltage limits, respectively.
[0008] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the process of inputting the distribution network operation status data and reference scheduling actions into a pre-established residual policy learning network model, and outputting the adjustment amount of the reference scheduling actions, as follows: In the current state and reference action As input, the output is the adjustment amount for the reference scheduling action, and is mapped to the physical feasible region through action upper and lower bound constraints, as follows: in, For the output function of the neural network; This is the motion adjustment amount output by the residual network; The original motion variables output by the actor network; Let the possible actions calculated based on the current state be an upper bound. The maximum feasible discharge power for the energy storage device. The maximum feasible active power of the distributed generator; To establish a lower bound for the possible actions calculated based on the current state, The maximum feasible charging power for energy storage devices.
[0009] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: considering the uncertainties of load and renewable energy output in the distributed opportunity constraint, and constructing an uncertainty set based on the predicted values and their covariance matrix, wherein the load uncertainty is represented by a set based on the first and second moments, and the renewable energy uncertainty is characterized by an ellipsoidal uncertainty set. The opportunity constraints of the BLU rod include substation capacity, voltage amplitude, and line capacity constraints.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the process of inputting the cumulative voltage violation cost into a pre-established hybrid distribution model, as follows: In the constrained Markov decision process, a hybrid risk constraint is constructed, and a conditional value-at-risk (VAT) index is introduced to characterize the tail risk of the cumulative voltage violation cost, as follows: in, This is the long-run cumulative cost function corresponding to the voltage constraint; This is the discount factor.
[0011] A hybrid distribution combining Bernoulli and exponential distributions is used to probabilistically model the long-term cumulative voltage violation cost. The Bernoulli distribution characterizes the case where constraints are fully satisfied, while the exponential distribution characterizes the voltage violation penalty term, as follows: in, The probability that the constraint is completely satisfied; The Dirac function at which the value is zero; The parameter is The exponential distribution of has a probability density function of . ; The parameters of the mixed distribution are determined by moment matching, as follows: in, This represents the average of the cumulative costs. Given the variance of cumulative cost, and Post-parameter and Uniquely determined by moment matching; The conditional value of risk is calculated based on a hybrid distribution and used as a risk metric for voltage safety constraints. The conditional value of risk is then calculated in a closed-form manner, as follows: in, It refers to the level of risk; Conditional Value at Risk (VaR) Dependence on State, Action, and Risk Level ,as follows: in, For auxiliary items; The decision-making strategy maximizes the long-term cumulative reward while satisfying the conditional value-at-risk threshold constraint, as follows: in, It is the long-term cumulative reward of the intelligent agent; Is it following the strategy? The trajectory; This is the conditional risk value threshold corresponding to the voltage constraint.
[0012] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the policy solution based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints is processed by a collaborative training method of policy network and multi-evaluation network, using a Lagrange training framework based on soft actors and commentators, and achieving policy exploration and stable optimization by maximizing the expected cumulative reward and introducing an entropy regularization term.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the Lagrange training framework based on soft actors and critics uses the Lagrange multiplier method to optimize the security constraints, and achieves joint training and collaborative optimization of the policy network and each evaluation network by adaptively adjusting the entropy weights and constraint penalty weights.
[0014] Secondly, in order to achieve the above objectives, this invention discloses a knowledge-embedded hybrid distributed risk assessment distribution network dispatching system, comprising: The data acquisition module is used to acquire distribution network operation status data, input the distribution network operation status data into a pre-established intraday rolling time-domain dynamic programming model of the distribution network, and output reference scheduling actions by introducing bibliometric rod chance constraints. The intraday rolling time-domain dynamic programming model of the distribution network is used to characterize the evolution of distribution network operation status and scheduling decision process under multiple time scales, and a constrained Markov decision process is constructed based on the dynamic programming model. The primary processing module is used to input the distribution network operation status data and reference scheduling actions into the pre-established residual strategy learning network model, output the adjustment amount of the reference scheduling actions, and adjust and optimize the reference scheduling actions based on the adjustment amount of the reference scheduling actions to obtain the adjusted scheduling actions. The secondary processing module is used to obtain the cumulative voltage violation cost, input the cumulative voltage violation cost into a pre-established hybrid distribution model, and perform closed-loop analytical calculation of the conditional risk value based on the hybrid distribution model to output the tail risk assessment result. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. The scheduling output module is used to solve the strategy based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints. It adopts a strategy network and risk assessment network collaborative training method to obtain the optimal scheduling strategy under the safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are synergistically optimized. The tail risk assessment results are used as the hybrid distributed risk constraints.
[0015] The beneficial effects of this invention are: This invention proposes, on the one hand, a residual policy architecture based on reference scheduling knowledge embedding with split-blob opportunistic constraints. This architecture introduces reference scheduling information obtained from split-blob optimization into the reinforcement learning process, enhancing the feasibility guidance capability and convergence efficiency of policy learning. On the other hand, it proposes a hybrid distributed risk modeling method to characterize tail risk characteristics that are difficult to express using traditional expectation-based risk metrics. This enables refined modeling and constraint expression of asymmetric risks such as voltage exceedances, thereby achieving coordinated optimization of economy and safety during real-time operation of the distribution network, considering source-load uncertainties. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2This is a schematic diagram of the residual policy network architecture based on the embedding of reference policy knowledge of the distributed bar opportunity constraint in this invention; Figure 3 This is a schematic diagram illustrating how the embedding of reference strategy knowledge improves exploration efficiency according to the present invention. Figure 4 This is a schematic diagram of the improved IEEE-33 node distribution network topology and distributed power source access in an embodiment of the present invention; Figure 5 This is a schematic diagram comparing the test performance of the embodiments of the present invention with other baseline methods; Figure 6 This is a schematic diagram comparing the fitting results of the voltage limit empirical distribution of the mixed distribution and the Gaussian distribution in an embodiment of the present invention; Figure 7 This is a schematic diagram comparing the safety performance of embodiments of the present invention with other methods; Figure 8 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1: like Figure 1 As shown, a hybrid distributed risk assessment and distribution network dispatching method based on knowledge embedding includes the following steps: S101: Obtain distribution network operation status data, input the distribution network operation status data into the pre-established intraday rolling time-domain dynamic programming model of the distribution network, and output the reference scheduling action by introducing the sub-Bruker chance constraint. The intraday rolling time-domain dynamic programming model of the distribution network is used to characterize the evolution of the distribution network operation status and the scheduling decision process under multiple time scales, and construct a constrained Markov decision process based on the dynamic programming model. The intraday rolling time-domain dynamic programming model for distribution networks extracts the state transition relationships at a single time point, describing the distribution network operation process as a constrained Markov decision process. It defines state variables, action variables, and a reward function based on the single-step operating cost. By describing the optimal scheduling process of the distribution network at multiple time scales, the scheduling scope is... ,as follows: in, This is a function of total operating cost; This is the cost function for a single-step operation. The power flow equation is as follows; For operational constraints; and In uncertainty The risk operators defined above can be instantiated as expected, conditional value of risk, or sub-Bruker risk measures according to scheduling requirements; These are decision variables, defined within the feasible region determined by the operational constraints. This includes the active power output of distributed generators and the charging and discharging power of energy storage devices; These are state variables, including power flow state variables such as voltage distribution and branch power flow.
[0019] State variables include active power of baseline load demand. and reactive power Active power of renewable energy units Remaining electricity of energy storage devices and the price of electricity transmission from substations ,as follows: in, For the state variables of the decision-making agent; The action variables are defined on nodes equipped with distributed generators and energy storage devices, as follows: in, A collection of nodes equipped with energy storage devices; Action variables generated for the decision-making agent; For the action variables corresponding to the output of the distributed generators at each node; Action variables corresponding to the charging and discharging power of energy storage devices at each node; The reward function based on the single-step running cost is as follows: Where, r t For agent reward variables; This represents the active power price of the power grid at time t; This indicates the active power transmitted from the upper-level power grid to the local power grid; This represents the active power quotation of the distributed generator connected to node n; This represents the active power output of the distributed generator connected to node n; This represents the set of nodes connected to distributed generators in the power grid; The cost function takes into account voltage safety costs, as follows: in, For voltage safety costs; To be a function that takes positive values; and These represent the minimum and maximum permissible voltage limits, respectively.
[0020] The uncertainty of load and renewable energy output is considered in the Bruker opportunity constraint, and an uncertainty set is constructed based on the predicted values and their covariance matrix. The load uncertainty is represented by a set based on the first and second moments, and the renewable energy uncertainty is characterized by an ellipsoidal uncertainty set. The opportunity constraints of the BLU rod include substation capacity, voltage amplitude, and line capacity constraints.
[0021] Specifically, the uncertainty of load and renewable energy output is considered in the Bruker chance constraint model, and an uncertainty set is constructed based on the predicted values and their covariance matrix. The load uncertainty is represented by a set based on the first and second moments, and the renewable energy uncertainty is characterized by an ellipsoidal uncertainty set, as follows: in, Let load power be a random variable. and Let be the load power vector at time t; For the first moment component of the load power random variable; For the second moment components of the load power random variable; Let the active power of renewable energy be a random variable. Let be the renewable energy power vector at time t; For the first-order moment component of the active power of renewable energy; For the second-order moment component of the active power of renewable energy; The node net load power is a random variable that combines load power and renewable energy power. and These are the first and second moments of the node's net load power, respectively; and These are the active power random variable and reactive power random variable of the net load at node t, respectively. and These are the first-order moment components of the active and reactive power of the node net load, respectively; , , and It represents the second-order moment components of the net active and reactive power of the node load.
[0022] The uncertainty set of node net load power is as follows: in, It is used to characterize random variables The set of Blue bars; express Defined by the constraint set on the right; For random variables The probability distribution it follows; For expectation operators; It is the inverse of the covariance matrix; The constraint parameter for the radius of the uncertain set of the ellipsoid; The constraint parameter is the semi-definite cone uncertainty set.
[0023] The constraints on substation capacity, voltage amplitude, and line capacity are reformulated as distributed robust chance constraints. After rewriting as distributed robust chance constraints, both the upper and lower bound constraints can be easily decomposed into two unilateral constraints, as follows: in, For the infimum operator, take the one with the smallest probability of satisfaction; It is a probability measure; It is a random variable; For decision variables; It is an uncertain set; It is a linear function of the decision variable; The coefficients are constant vectors; To constrain the upper limit; This represents the confidence level parameter.
[0024] Assuming uncertain variables Under the condition of following a normal distribution, It still follows a normal distribution, with a mean of Its variance is Based on uncertainty set The distributed robust chance constraint can be given a deterministic formulation, as follows: in, It is the inverse function of the cumulative distribution of the standard normal distribution; For the first moment components of the uncertain variable; Let be the second-order moment components of the uncertain variables. The split-Browl chance constraint is transformed into an equivalent deterministic linear inequality constraint, and further transformed into a mixed integer programming problem with a small number of 0-1 variables for solution.
[0025] S102: Input the distribution network operation status data and reference dispatch actions into the pre-established residual strategy learning network model, output the adjustment amount of the reference dispatch actions, and adjust and optimize the reference dispatch actions based on the adjustment amount of the reference dispatch actions to obtain the adjusted dispatch actions. The process of inputting distribution network operation status data and reference scheduling actions into a pre-established residual policy learning network model and outputting the adjustment amount of the reference scheduling actions is as follows: In the current state and reference action As input, the output is the adjustment amount for the reference scheduling action, and is mapped to the physical feasible region through action upper and lower bound constraints, as follows: in, For the output function of the neural network; This is the motion adjustment amount output by the residual network; The original motion variables output by the actor network; Let the possible actions calculated based on the current state be an upper bound. The maximum feasible discharge power for the energy storage device. The maximum feasible active power of the distributed generator; To establish a lower bound for the possible actions calculated based on the current state, The maximum feasible charging power for energy storage devices.
[0026] S103: Obtain the cumulative voltage violation cost, input the cumulative voltage violation cost into the pre-established hybrid distribution model, and perform closed-loop analytical calculation of the conditional value of risk based on the hybrid distribution model to output the tail risk assessment result. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. The process of inputting the cumulative voltage violation cost into the pre-established hybrid distribution model is as follows: In the constrained Markov decision process, a hybrid risk constraint is constructed, and a conditional value-at-risk (VAT) index is introduced to characterize the tail risk of the cumulative voltage violation cost, as follows: in, This is the long-run cumulative cost function corresponding to the voltage constraint; This is the discount factor.
[0027] A hybrid distribution combining Bernoulli and exponential distributions is used to probabilistically model the long-term cumulative voltage violation cost. The Bernoulli distribution characterizes the case where constraints are fully satisfied, while the exponential distribution characterizes the voltage violation penalty term, as follows: in, The probability that the constraint is completely satisfied; The Dirac function at which the value is zero; The parameter is The exponential distribution of has a probability density function of . .
[0028] The parameters of the mixed distribution are determined by moment matching, as follows: in, This represents the average of the cumulative costs. Let be the variance of the cumulative cost. Given... and Post-parameter and It can be uniquely determined through moment matching.
[0029] The conditional value of risk (VoV) is calculated using a hybrid distribution method and used as a risk metric for voltage safety constraints. The closed-form calculation of the VoV is as follows: in, It refers to the level of risk.
[0030] Conditional Value at Risk (VaR) Dependence on State, Action, and Risk Level ,as follows: in, This is an auxiliary item.
[0031] The final constructed constrained Markov decision process based on a hybrid distribution is characterized by an improvement on security risk. The decision strategy maximizes the long-term cumulative reward under the condition of satisfying the conditional risk value threshold constraint, as follows: in, It is the long-term cumulative reward of the intelligent agent; Is it following the strategy? The trajectory; This is the conditional risk value threshold corresponding to the voltage constraint.
[0032] S104: Based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints, the strategy is solved by co-training the strategy network and the risk assessment network to obtain the optimal scheduling strategy that meets the safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are optimized in a coordinated manner. The tail risk assessment results are used as the mixed distributed risk constraints.
[0033] Based on the Bellman equations, a state-value projection model is performed on the expected cumulative cost and the variance of cumulative cost, as follows: in, It is the cost function; This represents the state transition probability distribution; For an agent's policy, that is, in the state The probability of making the next action.
[0034] Construct loss functions for the evaluation agents based on the expected cumulative cost. Loss function for evaluating agents based on cumulative cost variance ,as follows: in, Evaluation agent for cumulative cost expectation Neural network parameters; Evaluation agent for cumulative cost variance Neural network parameters; Expected value Time difference target; variance Time difference target; A playback buffer for storing historical experience samples.
[0035] Based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints, the policy solution is processed by a collaborative training method of policy network and multi-evaluation network. A Lagrange training framework based on soft actors and critics is used to achieve policy exploration and stable optimization by maximizing the expected cumulative reward and introducing an entropy regularization term.
[0036] Based on the Lagrange training framework of the soft actor-commentator, policy exploration and stable optimization are achieved by maximizing the expected cumulative reward and introducing an entropy regularization term, with the following objectives: in, It is an adaptive entropy weight; It is an adaptive penalty weight; It is a strategy Entropy; It is a given entropy threshold used to constrain the exploration level of the strategy to ensure the minimum exploration requirement.
[0037] A strategy network, a double-Q value network, a cumulative cost expectation evaluation network, and a cumulative cost variance evaluation network are constructed to estimate strategy return and risk-related indicators, respectively. (Decision-making agent) The neural network parameters are The cumulative reward evaluation agent uses a dual-Q design. The neural network parameters are , The neural network parameters are Cumulative cost expected assessment agent The neural network parameters are Cumulative cost variance evaluation agent The neural network parameters are .
[0038] The Lagrange training framework based on soft actors and commentators uses the Lagrange multiplier method to optimize safety constraints. By adaptively adjusting the entropy weights and constraint penalty weights, it achieves joint training and collaborative optimization of the policy network and various evaluation networks.
[0039] Specifically, the Lagrange multiplier method is used to optimize the security constraints. By adaptively adjusting the entropy weights and constraint penalty weights, joint training and collaborative optimization of the policy network and each evaluation network are achieved.
[0040] Loss function for design decision-making agent and the loss function for evaluating the agent based on cumulative rewards ,as follows: in, The target reward is calculated using the Bellman equation.
[0041] The constraints are optimized using the Lagrange multiplier method, and the weights are adaptively adjusted through learning. and To minimize entropy loss and cost loss, as follows: in, This is the loss function corresponding to the temperature entropy weights; This is the loss function corresponding to the weights of the penalty terms.
[0042] Specifically, the following embodiments further illustrate the solution of the present invention: This embodiment proposes a residual policy network architecture based on the embedding of reference policy knowledge under the split-brow chance constraint (e.g., Figure 2 (As shown). The core idea is that the policy network does not directly output the final scheduling decision. Instead, it takes the current system state and the reference action generated by the optimization model as input, and outputs the residual correction amount for the reference action to obtain the final scheduling action. The reference action originates from a distributed random chance constraint optimization model that explicitly considers system uncertainties, thus embedding physical constraints and operational knowledge into the reinforcement learning decision-making process. Compared to purely data-driven spatiotemporal graph feature extraction methods, this framework exhibits stronger robustness when facing operational scenarios outside the training distribution. This method uses the reference action as the policy initialization benchmark, allowing the policy to naturally explore operation points with physical feasibility and economic significance, thereby significantly improving the policy exploration efficiency. Simultaneously, by transforming the learning objective from direct policy generation to residual correction learning, the learning complexity of the policy network is effectively reduced, making the training process more stable. For system resources at time t, the upper and lower bounds of the corresponding actionable actions are calculated based on their current operating state to ensure that the output action always satisfies the physical feasibility constraints. Figure 3 This study demonstrates the differences in exploration behavior under different constraint scenarios, with and without reference policy guidance. In scenarios with no (or weak) constraints, although a global optimum exists, policies without knowledge embedding tend to converge to suboptimal economic regions due to initialization randomness and insufficient exploration efficiency. When the feasible region exhibits a fragmented structure, the exploration process of methods without knowledge embedding is more easily restricted by safety constraints, further reducing the probability of reaching the global optimum. In contrast, policies guided by reference actions converge to the global optimum more stably because they are initialized near high-quality feasible regions.
[0043] This implementation method uses an improved IEEE-33 node distribution network system for testing, with distributed energy resources and access nodes such as... Figure 4 As shown. The experimental data uses two years of half-hourly load, photovoltaic output, and electricity price data, with 700 days used for training and 30 days for testing. A decision-making agent was constructed and trained according to the method described in this patent. During training, the agent was run on a 30-day test set every 1000 training rounds to verify the average daily system operating cost. A soft actor-critic algorithm without knowledge embedding, a soft actor-critic algorithm with knowledge embedding based on expectation constraints, and a soft actor-critic algorithm with knowledge embedding based on Gaussian distribution constraints were selected as comparison methods, and their performance was evaluated on the same test set. The comparison results are shown below. Figure 5As shown in Table 1. The results show that the method of the present invention achieves the best performance in terms of cumulative reward, while the method without knowledge embedding is prone to economic suboptimality due to its lower exploration efficiency. Furthermore, statistical analysis was performed on the distribution of voltage violation experience collected during the training process, such as... Figure 6 As shown. The results indicate that the voltage violation cost distribution exhibits significant heavy-tailed and skewed characteristics, with a large number of samples concentrated in the zero or low violation region, but a small number of extreme events dominate the tail risk. Traditional Gaussian distribution models struggle to characterize this type of asymmetric distribution structure, especially exhibiting significant underestimation in tail probability estimation. In contrast, the hybrid distribution model proposed in this invention can more accurately fit the overall distribution characteristics, especially demonstrating higher fitting accuracy in the tail region, thus providing a more reliable statistical basis for risk-sensitive decision-making. In the safety comparison analysis, since the knowledge-free embedding method cannot converge stably, the safety performance comparison mainly focuses on the following four methods: expectation-based knowledge embedding method, knowledge embedding method based on hybrid distribution (… ), knowledge embedding method based on Gaussian distribution ( ), Knowledge embedding method based on hybrid distribution ( ).like Figure 7 As shown, statistical analysis was performed on the results of two months of operation. The left y-axis of the figure is a box plot distribution of violation costs, where y represents the [5, 95] percentile. The right y-axis of the figure is the unsafe day rate, defined as the percentage of days with constraint violations, used to directly measure the safety of system operation. The results show that the expectation-based risk modeling method has the worst safety, with an unsafe day rate exceeding 50%; the mixed distribution ( The method achieved similar average performance; based on Gaussian distribution ( While the method based on the mixed distribution reduced the unsafe day rate to below 20%, it also resulted in rare and large-amplitude voltage violations. This further illustrates that the Gaussian distribution is ineffective in representing the long-tailed characteristics of actual risks and exhibits a significant distribution mismatch problem. The method exhibits the best robustness, with the unsafe day rate reduced to below 1%, the most concentrated distribution of violation costs, and a significant reduction in tail risk.
[0044] Table 1 Comparison of test performance results of the embodiments of the present invention with other baseline methods Example 2: To achieve the above objective, such as Figure 8 As shown, based on Embodiment 1, this invention discloses a knowledge-embedded hybrid distributed risk assessment distribution network dispatching system, comprising: The data acquisition module 11 is used to acquire the distribution network operation status data, input the distribution network operation status data into the pre-established intraday rolling time-domain dynamic programming model of the distribution network, and output the reference scheduling action by introducing the sub-Bruker chance constraint. The intraday rolling time-domain dynamic programming model of the distribution network is used to characterize the evolution of the distribution network operation status and the scheduling decision process under multiple time scales, and constructs a constrained Markov decision process based on the dynamic programming model. The primary processing module 12 is used to input the distribution network operation status data and reference scheduling actions into the pre-established residual strategy learning network model, output the adjustment amount of the reference scheduling actions, and adjust and optimize the reference scheduling actions based on the adjustment amount of the reference scheduling actions to obtain the adjusted scheduling actions. The secondary processing module 13 is used to obtain the cumulative voltage violation cost, input the cumulative voltage violation cost into the pre-established hybrid distribution model, and perform closed-loop analytical calculation of the conditional risk value based on the hybrid distribution model to output the tail risk assessment result. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. The scheduling output module 14 is used to solve the strategy based on the tail risk assessment results, the adjusted scheduling actions and the preset conditional risk value constraints, by adopting a strategy network and risk assessment network collaborative training method to obtain the optimal scheduling strategy under the condition of safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are synergistically optimized. The tail risk assessment results are used as the hybrid distributed risk constraints.
[0045] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.
[0046] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0047] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0048] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.
Claims
1. A hybrid distributed risk assessment and distribution network dispatching method based on knowledge embedding, characterized in that, The method includes the following steps: The distribution network operation status data is acquired and input into a pre-established intraday rolling time-domain dynamic programming model. By introducing bibliometric chance constraints, a reference dispatch action is output. The intraday rolling time-domain dynamic programming model is used to characterize the evolution of the distribution network operation status and the dispatch decision process under multiple time scales, and a constrained Markov decision process is constructed based on the dynamic programming model. The distribution network operation status data and reference scheduling actions are input into a pre-established residual strategy learning network model, and the adjustment amount of the reference scheduling actions is output. Based on the adjustment amount of the reference scheduling actions, the reference scheduling actions are adjusted and optimized to obtain the adjusted scheduling actions. The cumulative voltage violation cost is obtained, and then input into a pre-established hybrid distribution model. Based on the hybrid distribution model, the conditional value of risk is calculated using closed-loop analytical methods, and the tail risk assessment result is output. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. Based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints, the strategy is solved by co-training the strategy network and the risk assessment network to obtain the optimal scheduling strategy that meets the safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are optimized in a coordinated manner. The tail risk assessment results are used as a hybrid distributed risk constraint.
2. The knowledge-embedded hybrid distributed risk assessment and distribution network dispatching method according to claim 1, characterized in that, The intraday rolling time-domain dynamic programming model for the distribution network extracts the state transition relationships at a single time segment, describing the distribution network operation process as a constrained Markov decision process, and defines state variables, action variables, and a reward function based on the single-step operation cost. By describing the optimal scheduling process of the distribution network at multiple time scales, the scheduling scope is... ,as follows: in, This is a function of total operating cost; This is the cost function for a single-step operation. The power flow equation is as follows; For operational constraints; and In uncertainty Risk operators defined above; These are decision variables, defined within the feasible region determined by the operational constraints. This includes the active power output of distributed generators and the charging and discharging power of energy storage devices; It is a state variable.
3. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 2, characterized in that, The state variables include active power of baseline load demand and reactive power , active power of renewable energy units , remaining power of energy storage devices and substation power transmission price as follows: wherein, is a state variable of the decision-making agent; The action variables are defined on nodes equipped with distributed generators and energy storage devices, as follows: in, A collection of nodes equipped with energy storage devices; Action variables generated for the decision-making agent; For the action variables corresponding to the output of the distributed generators at each node; Action variables corresponding to the charging and discharging power of energy storage devices at each node; The reward function based on the single-step running cost is as follows: Where, r t For agent reward variables; This represents the active power price of the power grid at time t; This indicates the active power transmitted from the upper-level power grid to the local power grid; This represents the active power quotation of the distributed generator connected to node n; This represents the active power output of the distributed generator connected to node n; This represents the set of nodes connected to distributed generators in the power grid; The cost function takes into account voltage safety costs, as follows: in, For voltage safety costs; To be a function that takes positive values; and These represent the minimum and maximum permissible voltage limits, respectively.
4. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 1, characterized in that, The process of inputting distribution network operation status data and reference scheduling actions into a pre-established residual strategy learning network model and outputting the adjustment amount of the reference scheduling actions is as follows: In the current state and reference action As input, the output is the adjustment amount for the reference scheduling action, and is mapped to the physical feasible region through action upper and lower bound constraints, as follows: in, For the output function of the neural network; This is the motion adjustment amount output by the residual network; The original motion variables output by the actor network; Let the possible actions calculated based on the current state be an upper bound. The maximum feasible discharge power for the energy storage device. The maximum feasible active power of the distributed generator; To establish a lower bound for the possible actions calculated based on the current state, The maximum feasible charging power for energy storage devices.
5. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 1, characterized in that, The proposed multi-bar opportunity constraint considers the uncertainties of load and renewable energy output, and constructs an uncertainty set based on the predicted values and their covariance matrix. Load uncertainty is represented using a set based on first and second moments, while renewable energy uncertainty is characterized using an ellipsoidal uncertainty set. The opportunity constraints of the BLU rod include substation capacity, voltage amplitude, and line capacity constraints.
6. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 1, wherein, The process of inputting the cumulative voltage violation cost into the pre-established hybrid distribution model is as follows: In the constrained Markov decision process, a hybrid risk constraint is constructed, and a conditional value-at-risk (VAT) index is introduced to characterize the tail risk of the cumulative voltage violation cost, as follows: wherein, is a long-term accumulated cost function corresponding to the voltage constraint; is a discount factor; A hybrid distribution combining Bernoulli and exponential distributions is used to probabilistically model the long-term cumulative voltage violation cost. The Bernoulli distribution characterizes the case where constraints are fully satisfied, while the exponential distribution characterizes the voltage violation penalty term, as follows: in, The probability that the constraint is completely satisfied; The Dirac function at which the value is zero; The parameter is The exponential distribution of has a probability density function of . ; The parameters of the mixed distribution are determined by moment matching, as follows: wherein, is the mean of the cumulative cost; is the variance of the cumulative cost, given and the posterior parameters and are uniquely determined by the moment matching. The conditional value of risk is calculated based on a hybrid distribution and used as a risk metric for voltage safety constraints. The conditional value of risk is then calculated in a closed-form manner, as follows: in, It refers to the level of risk; Conditional value at risk dependence on state, action, and risk level As follows: wherein is an auxiliary term; The decision-making strategy maximizes the long-term cumulative reward while satisfying the conditional value-at-risk threshold constraint, as follows: wherein, is the long-term cumulative reward of the agent; is a trajectory that follows the policy ; is the conditional value-at-risk threshold corresponding to the voltage constraint.
7. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 1, characterized in that, The strategy solution based on tail risk assessment results, adjusted scheduling actions, and preset conditional risk value constraints uses a Lagrange training framework based on soft actors and commentators. It achieves strategy exploration and stable optimization by maximizing expected cumulative reward and introducing an entropy regularization term.
8. The knowledge embedding based hybrid distribution risk assessment power grid dispatching method according to claim 7, characterized in that, The Lagrange training framework based on soft actors and commentators uses the Lagrange multiplier method to optimize security constraints. By adaptively adjusting the entropy weights and constraint penalty weights, it achieves joint training and collaborative optimization of the policy network and each evaluation network.
9. A knowledge-embedded hybrid distributed risk assessment distribution network dispatching system, employing the knowledge-embedded hybrid distributed risk assessment distribution network dispatching method as described in any one of claims 1 to 8, characterized in that... include: The data acquisition module is used to acquire distribution network operation status data, input the distribution network operation status data into a pre-established intraday rolling time-domain dynamic programming model of the distribution network, and output reference scheduling actions by introducing bibliometric rod chance constraints. The intraday rolling time-domain dynamic programming model of the distribution network is used to characterize the evolution of distribution network operation status and scheduling decision process under multiple time scales, and a constrained Markov decision process is constructed based on the dynamic programming model. The primary processing module is used to input the distribution network operation status data and reference scheduling actions into the pre-established residual strategy learning network model, output the adjustment amount of the reference scheduling actions, and adjust and optimize the reference scheduling actions based on the adjustment amount of the reference scheduling actions to obtain the adjusted scheduling actions. The secondary processing module is used to obtain the cumulative voltage violation cost, input the cumulative voltage violation cost into a pre-established hybrid distribution model, and perform closed-loop analytical calculation of the conditional risk value based on the hybrid distribution model to output the tail risk assessment result. The hybrid distribution model is used to characterize the long-tail probability distribution of the cumulative voltage violation cost. The scheduling output module is used to solve the strategy based on the tail risk assessment results, the adjusted scheduling actions, and the preset conditional risk value constraints. It adopts a strategy network and risk assessment network collaborative training method to obtain the optimal scheduling strategy under the safety risk constraints. Based on the optimal scheduling strategy, the operating cost and safety of the distribution network are synergistically optimized. The tail risk assessment results are used as the hybrid distributed risk constraints.
10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it employs the knowledge embedding-based hybrid distributed risk assessment distribution network scheduling method as described in any one of claims 1 to 8.