Energy and power hierarchical probability balance control method and system based on distribution network grid

By constructing a hierarchical probabilistic balance control method for distribution grids, combined with reinforcement learning and multi-scale optimization, the problems of energy-power dynamic balance inaccuracy and poor cross-layer control coupling caused by the access of new energy to the distribution network are solved, and the absorption of a high proportion of new energy and the improvement of system robustness are achieved.

CN120675110APending Publication Date: 2025-09-19ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510837524.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies are unable to cope with the dual uncertainties brought about by errors in renewable energy output forecasts and fluctuations across time scales, resulting in inaccurate planning of inter-regional power interaction capacity, delayed response of energy storage charging and discharging strategies and load regulation, and easily causing local power exceeding limits or insufficient backup capacity, affecting the safe and economical operation of the system. In addition, there is a lack of a hierarchical coordination mechanism, making it difficult to strike a balance between global optimization and unit-level refined control.

Method used

A hierarchical probabilistic balance control method for energy and power based on the distribution grid is adopted. Through hierarchical probabilistic constraint modeling, reinforcement learning dynamic gain adjustment and multi-scale collaborative optimization, a three-level control architecture of region-grid-unit is constructed. The confidence scenario set is constructed by combining QLSTM time series prediction, SSA decomposition correction and Beta-Weibull distribution. The gain control instructions are generated using the Actor-Critic reinforcement learning framework, and iterative optimization is performed through ADMM collaborative optimization to coordinate cross-layer power flow, energy storage and load resources.

Benefits of technology

It effectively solves the problems of energy-power dynamic balance imbalance and poor cross-layer control coupling caused by the high proportion of new energy access to the distribution network, improves the distribution network's ability to absorb new energy and operational robustness, and realizes refined control from the regional level to the unit level, ensuring safe and economical operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675110A_ABST
    Figure CN120675110A_ABST
Patent Text Reader

Abstract

The invention discloses an energy and power hierarchical probability balance control method and system based on a distribution network grid, and relates to the technical field of power grid distribution network control, and the method comprises the following steps: predicting new energy output based on a quantum enhancement prediction algorithm, and constructing a confidence scene set in combination with a pre-obtained distribution model; establishing a region-grid-unit three-level control architecture, and obtaining control parameters based on the confidence scene set and a preset hierarchical probability constraint; according to the control parameters and the power grid state data, by taking energy and power balance as a target, generating a gain control instruction through a strategy-value collaborative reinforcement learning method; performing a plurality of rounds of iterative optimization on the gain control instruction based on a distributed collaborative optimization mechanism; according to the optimized gain control instruction, each layer of execution unit is controlled, and power flow, energy storage and load resources are coordinated. The method is used for solving the problems of misalignment of energy-power dynamic balance and poor cross-layer control coupling caused by access of high-proportion new energy to the power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power distribution network control, and more specifically, to a method and system for controlling energy and power hierarchical probability balance based on a distribution network. Background Art

[0002] Probabilistic balancing of grid energy and power is a core issue in ensuring the safe and stable operation of power systems. Its essence is to achieve dynamic matching between supply and demand under uncertain conditions. Factors such as fluctuations in renewable energy output and random variations in load demand in power systems pose challenges to maintaining a real-time balance between active and reactive power. Probabilistic balancing methods, incorporating probability theory and stochastic optimization theory, model uncertain variables such as renewable energy output and load probabilistically. Based on this model, robust control and dispatch strategies are developed to ensure power balance in the system under most scenarios, avoiding problems such as frequency deviation and voltage overshooting caused by extreme fluctuations.

[0003] Reinforcement learning is a machine learning paradigm in which an intelligent agent learns to take actions to maximize cumulative rewards through interaction with its environment. It does not rely on large amounts of labeled data, but instead optimizes its policy through self-learning. The learning process involves the agent iterating its policy through trial and error, balancing exploration and exploitation. During this process, the agent adjusts its policy based on reward signals from the environment to achieve optimal behavior. Reinforcement learning has extensive and important applications in the control field. In industrial control, reinforcement learning can be used to optimize production processes and equipment maintenance, improving the efficiency and quality of industrial production. Furthermore, in robust control, reinforcement learning learns and optimizes policies through autonomous exploration of the environment to handle system uncertainties and enhance the system's adaptability and robustness to interference.

[0004] For example, the invention patent announcement with announcement number: CN114371619B discloses an MGT-CCHP variable operating condition dynamic energy efficiency optimization control method, which involves establishing an MGT-CCHP input and output training set, inputting the training set into a neural network, and training through an error back propagation algorithm to obtain a nonlinear prediction model; obtaining a dynamic energy efficiency index, a load tracking optimization target, and a control quantity change optimization target, and reconstructing the first control objective function through a Utopia point tracking control framework based on the above three optimization targets; establishing a second control objective function, and setting the switching conditions between the first control objective function and the second control objective function; obtaining an optimal control increment sequence based on the nonlinear prediction model, the first control objective function, and the second control objective function, in combination with the input quantity at the current moment and the input quantity at the previous moment; obtaining the control input increment at the current moment based on the optimal control increment sequence, and calculating the control quantity at the current moment in combination with the control quantity at the previous moment, and inputting the control quantity at the current moment into the MGT-CCHP.

[0005] For example, the invention patent announcement with announcement number: CN116699979A discloses an optimal control method for a distributed multi-agent system based on policy gradient, which relates to an optimal control method for a distributed multi-agent system based on policy gradient, and is characterized in that it includes the following steps: constructing a corresponding state-action value function based on the performance indicators of the discrete-time nonlinear multi-agent system; determining an initial control strategy within the allowable control set; performing strategy evaluation based on multiple state-action value functions and the current control strategy; improving the strategy based on the strategy evaluation result in combination with the gradient descent method to obtain the strategy for the next round of iteration, returning to perform strategy evaluation until a converged optimal control strategy is learned, wherein the iterative optimization process is implemented based on the Actor-Critic structure and the experience replay mechanism, adopting the critic network to approximate the control strategy, and adopting the actor network to approximate the state-action value function; and controlling the distributed multi-agent system based on the optimal control strategy.

[0006] The above disclosed technical solutions have at least the following technical problems: Existing technologies struggle to address the dual uncertainties of renewable energy output forecast errors and cross-timescale fluctuations. This leads to inaccurate capacity planning for inter-regional power interaction, delayed energy storage charging and discharging strategies, and load regulation responses. This can easily lead to local power overruns or insufficient backup capacity, impacting the safe and economic operation of the system. Furthermore, existing control architectures lack a layered coordination mechanism, making it difficult to strike a balance between global optimization and refined unit-level control. This makes it impossible to effectively coordinate the multi-dimensional constraints of cross-layer power transmission, dynamic energy storage response, and flexible load regulation.

[0007] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0008] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an energy and power hierarchical probabilistic balance control method and system based on a distribution grid, which solves the problems of energy-power dynamic balance inaccuracy and poor cross-layer control coupling caused by a high proportion of new energy access to the distribution network through hierarchical probabilistic constraint modeling, reinforcement learning dynamic gain adjustment and multi-scale collaborative optimization.

[0009] To achieve the above object, the present invention provides the following technical solutions: The energy and power hierarchical probabilistic balance control method based on the distribution grid includes the following steps: predicting the output of renewable energy based on the quantum enhanced prediction algorithm, and building a confidence scenario set in combination with a pre-acquired distribution model; establishing a three-level control architecture of region-grid-unit, and obtaining control parameters based on the confidence scenario set and preset hierarchical probabilistic constraints; generating gain control instructions based on the control parameters and grid status data, with energy and power balance as the goal, through a strategy-value collaborative reinforcement learning method; performing several rounds of iterative optimization on the gain control instructions based on a distributed collaborative optimization mechanism; and controlling the execution units at each layer according to the optimized gain control instructions to coordinate power flow, energy storage, and load resources.

[0010] In a preferred embodiment, the control parameters include cross-layer power transmission capacity, backup coefficient and self-sufficiency parameter; the grid status data includes energy residual, power residual and energy storage data; the gain control instructions include energy gain coefficient and power gain coefficient of each layer.

[0011] In a preferred embodiment, the method for constructing the confidence scenario set is specifically as follows: based on SSA, decompose the pre-acquired output time series sequence, and obtain several groups of principal component subsequences according to singular value grouping, convert the principal component subsequences back into time series sequences to obtain a corrected output time series sequence; use the corrected output time series sequence as training data, train the pre-acquired QLSTM output prediction model and perform prediction to obtain an output prediction sequence, wherein the output includes photovoltaic output and wind power output; according to the output prediction sequence, based on the photovoltaic output obeying the Beta distribution and the wind power output obeying the Weibull distribution, obtain a new energy output probability scenario set through Monte Carlo sampling; impose 1-norm constraints and ∞-norm constraints on the probability scenario set, screen typical scenarios through the K-means clustering algorithm, and output a new energy output confidence scenario set.

[0012] In a preferred embodiment, the method for obtaining the control parameters is specifically as follows: according to the confidence scenario set and regional layer constraints, with the internal power balance of each region as the goal, solving the cross-layer power transmission capacity; presetting the backup coefficient and determining the grid layer constraints; obtaining the distributed new energy capacity and energy storage capacity of the unit layer, and based on the self-sufficiency capacity constraints of the unit layer, determining the self-sufficiency parameters through pre-set confidence intervals.

[0013] In a preferred embodiment, the method for obtaining the gain control instruction is specifically as follows: mapping the control parameters and grid state data into a standardized state vector through a state encoder; based on the standardized state vector, based on the PPO algorithm, outputting the energy gain coefficient and power gain coefficient of each layer through the Actor network, and mapping the gain coefficient to the [-1, 1] interval through the tanh function; based on the mapped gain coefficient, calculating the action value through the Critic network, and combining with the preset reward function, updating the Actor and Critic networks through the policy gradient, and outputting the gain control instruction.

[0014] In a preferred embodiment, the gain control instruction is subjected to several rounds of iterative optimization, specifically: based on the ADMM collaborative mechanism, an augmented Lagrangian function is constructed in a digital twin environment, the control parameters are used as control variables and the multipliers are initialized; according to the energy residual and the power residual, the control variables and multipliers of each layer are updated alternately, and the maximum number of iterations is preset until the residual converges or the maximum number of iterations is reached, and the tensor of the optimized multiplier and the optimized control parameters are obtained; the grid state data, the tensor of the optimized multiplier and the optimized control parameters are used as the state at the next moment, and the gain control instruction is updated.

[0015] In a preferred embodiment, the control of each layer execution unit coordinates power flow, energy storage and load resources, specifically: controlling the cross-layer power transmission capacity according to the optimized gain control instruction; and controlling the energy storage charging power and energy storage discharging power of the grid layer based on a preset energy storage dynamic model; and controlling the adjustable load power of the unit layer.

[0016] In a preferred embodiment, the updating of the Actor and Critic networks is specifically as follows: updating the Actor network by back propagation with the goal of minimizing energy and power imbalance; and updating the Critic network with the goal of minimizing the temporal difference error of the Q value.

[0017] The energy and power hierarchical probabilistic balance control system based on the distribution grid includes: a scenario set construction module, which is used to predict the output of renewable energy based on the quantum enhanced prediction algorithm, and build a confidence scenario set in combination with a pre-acquired distribution model; a control parameter acquisition module, which is used to establish a three-level control architecture of region-grid-unit, and obtain control parameters based on the confidence scenario set and preset hierarchical probabilistic constraints; a gain control instruction acquisition module, which is used to generate gain control instructions based on the control parameters and grid status data, with energy and power balance as the goal, through a strategy-value collaborative reinforcement learning method; a gain control instruction update module, which is used to perform several rounds of iterative optimization of the gain control instructions based on a distributed collaborative optimization mechanism; and a hierarchical control module, which is used to control the execution units at each layer according to the optimized gain control instructions, and coordinate power flow, energy storage and load resources.

[0018] The technical effects and advantages of the energy and power layered probability balance control method and system based on the distribution grid of the present invention are as follows: This invention combines QLSTM time series prediction, SSA decomposition correction, and Beta-Weibull distribution to construct a set of confidence scenarios for renewable energy output. Based on this, it innovatively employs hierarchical probabilistic constraint modeling and an actor-critic reinforcement learning framework to achieve dynamic balance in both energy and power dimensions. Through a hierarchical architecture (regional, grid, and unit levels), distribution network control granularity is defined. Reinforcement learning is used to generate gain control commands, and multiple rounds of iterative corrections are performed in conjunction with ADMM collaborative optimization. Ultimately, this achieves cross-layer power interaction, global coordinated control of energy storage charging and discharging, and loads. Its core advantages lie in: dynamic gain adjustment based on reinforcement learning effectively addresses renewable energy fluctuations; the hierarchical probabilistic constraint mechanism balances system safety and operational economy; and multi-scale optimization enables refined control from the regional to the unit level. This effectively addresses the issues of energy-power dynamic imbalance and poor cross-layer control coupling caused by high renewable energy penetration in the distribution network, while significantly improving the distribution network's capacity to accommodate high renewable energy penetration and its operational robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic flow chart of a method for controlling energy and power stratified probability balance based on a distribution grid provided in an embodiment of the present invention.

[0020] Figure 2 A schematic diagram of the structure of a distribution grid-based energy and power hierarchical probability balance control system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] Example 1, Figure 1 The present invention provides a method for controlling energy and power hierarchical probability balance based on a distribution grid, comprising the following steps: S1, predicts the output of new energy based on the quantum-enhanced prediction algorithm, combines it with the pre-acquired distribution model, and builds a confidence scenario set; S2, establishes a three-level control architecture of region-grid-cell, and obtains control parameters based on a set of confidence scenarios and preset hierarchical probability constraints; S3, based on the control parameters and grid status data, generates gain control instructions through a policy-value collaborative reinforcement learning method with the goal of energy and power balance; S4, performing several rounds of iterative optimization on the gain control instructions based on the distributed collaborative optimization mechanism; S5, according to the optimized gain control instructions, controls the execution units at each layer and coordinates the power flow, energy storage and load resources.

[0023] This embodiment combines QLSTM time series prediction, SSA decomposition correction, and the Beta-Weibull distribution to construct a set of renewable energy output confidence scenarios. Based on this, it innovatively employs hierarchical probabilistic constraint modeling and an actor-critic reinforcement learning framework to achieve dynamic balance in both energy and power dimensions. Through a hierarchical architecture (regional, grid, and unit levels), distribution network control granularity is defined. Reinforcement learning is used to generate gain control commands, and multiple rounds of iterative corrections are performed in conjunction with ADMM collaborative optimization. Ultimately, this achieves cross-layer power interaction and globally coordinated control of energy storage charging and discharging and loads. Its core advantages include: dynamic gain adjustment based on reinforcement learning effectively addresses renewable energy fluctuations; the hierarchical probabilistic constraint mechanism balances system safety and operational economy; and multi-scale optimization enables refined control from the regional to the unit level. This effectively addresses the issues of energy-power dynamic balance imbalance and poor cross-layer control coupling caused by high renewable energy penetration in the distribution network, significantly improving the distribution network's capacity to accommodate high renewable energy penetration and its operational robustness.

[0024] S1, predicts the output of new energy based on the quantum enhanced prediction algorithm, and builds a confidence scenario set in combination with the pre-acquired distribution model.

[0025] In this embodiment, the pre-acquired distribution models include a Beta distribution model and a Weibull distribution model.

[0026] In this embodiment, the method for constructing the trusted scene set is specifically as follows: Based on SSA, the pre-obtained output time series is decomposed and grouped according to singular values ​​to obtain several groups of principal component subsequences. The principal component subsequences are then converted back into time series to obtain the corrected output time series. The corrected output time series is used as training data to train a pre-acquired QLSTM output prediction model and perform predictions to obtain an output prediction sequence, where the output includes photovoltaic output and wind power output; Based on the output forecast sequence, and the assumption that photovoltaic output follows the Beta distribution and wind power output follows the Weibull distribution, a set of renewable energy output probability scenarios is obtained through Monte Carlo sampling. The 1-norm constraint and ∞-norm constraint are imposed on the probability scenario set, and the typical scenarios are screened through the K-means clustering algorithm to output the new energy output confidence scenario set.

[0027] In this embodiment, the specific formula of the confidence scene set is:

[0028] Where, is the set of confident scenes, is the total number of scenarios in the probability scenario set, is the probability value of the kth scene, is the baseline probability of the kth scene in the obtained probability scene set, is the preset 1-norm constraint parameter, is the preset ∞-norm constraint parameter, is the maximum value of the probability norm in k scenarios.

[0029] S2, establishes a three-level control architecture of region-grid-cell, and obtains control parameters based on a set of confidence scenarios and preset hierarchical probability constraints.

[0030] In this embodiment, the control parameters include cross-layer power transmission capacity, reserve coefficient and self-sufficiency parameter.

[0031] In this embodiment, the layered probability constraints include regional layer constraints, grid layer constraints, and unit layer self-sufficiency constraints.

[0032] In this embodiment, the method for obtaining the control parameters is specifically as follows: Based on the confidence scenario set and regional layer constraints, the cross-layer power transfer capacity is solved with the goal of internal power balance in each region. Preset backup coefficients and determine grid layer constraints; Obtain the distributed new energy capacity and energy storage capacity at the unit level, and determine the self-sufficiency parameters through pre-set confidence intervals based on the self-sufficiency constraints of the unit level.

[0033] In this embodiment, the regional layer constraint is specifically formulated as follows:

[0034] Where, is the number of grid layers in the regional layer, is the transmission capacity of the kth grid to the regional layer, is the total transmission capacity of the regional layer to the transmission grid, is the transmission capacity margin, is the confidence level of the regional layer.

[0035] In this embodiment, the grid layer constraint is specifically formulated as follows:

[0036] Where, is the number of unit layers in the grid layer, is the equilibrium ratio of the lth unit layer, is the maximum load of the lth unit layer, is the transmission capacity from the grid layer to the regional layer, To store energy in the grid layer, is the reserve coefficient, is the maximum total load of the unit layer in the grid layer, is the confidence level of the grid layer.

[0037] In this embodiment, the self-sufficiency constraint is specifically formulated as follows:

[0038] Where, To contribute to the new energy in the unit layer, For the energy storage output in the unit layer, is a self-sufficient parameter, is the maximum load in the unit layer, is the confidence level of the unit layer.

[0039] S3, based on the control parameters and grid status data, with the goal of energy and power balance, generates gain control instructions through a strategy-value collaborative reinforcement learning method.

[0040] In this embodiment, the grid state data includes energy residual, power residual and energy storage data; the gain control instruction includes the energy gain coefficient and power gain coefficient of each layer.

[0041] In this embodiment, the method for obtaining the gain control instruction is specifically as follows: Mapping control parameters and grid state data into a standardized state vector through a state encoder; According to the standardized state vector, based on the PPO algorithm, the energy gain coefficient and power gain coefficient of each layer are output through the Actor network, and the gain coefficient is mapped to the [-1, 1] interval through the tanh function; According to the mapped gain coefficient, the action value is calculated through the Critic network, and combined with the preset reward function, the Actor and Critic networks are updated through the policy gradient to output the gain control instruction.

[0042] In this embodiment, the updating of the Actor and Critic networks is specifically as follows: Update the Actor network through backpropagation with the goal of minimizing energy and power imbalance; The Critic network is updated with the goal of minimizing the temporal difference error of the Q value.

[0043] In this embodiment, the preset reward function has the following specific formula:

[0044] Where, is the reward function, is the energy residual of the Lth layer, is the power residual of the Lth layer, is the penalty weight of energy residual, is the penalty weight of power residual, is the penalty coefficient for the default probability of the Lth layer, is the probability of power imbalance occurring at the Lth layer, where L is the sequence number of the regional layer, grid layer, and unit layer.

[0045] In this embodiment, the gain control instruction is specifically formulated as follows:

[0046] Where, is the control action set of the gain control instruction, 、 and is the energy gain coefficient of each layer, 、 and is the power gain coefficient of each layer.

[0047] S4, performs several rounds of iterative optimization on the gain control instructions based on the distributed collaborative optimization mechanism.

[0048] In this embodiment, the gain control instruction is optimized for several rounds of iterations, specifically: Based on the ADMM collaborative mechanism, the augmented Lagrangian function is constructed in the digital twin environment, the control parameters are used as control variables and the multipliers are initialized; According to the energy residual and power residual, the control variables and multipliers of each layer are updated alternately, and the maximum number of iterations is preset until the residual converges or the maximum number of iterations is reached, and the tensor of the optimized multipliers and the optimized control parameters are obtained; The grid state data, the optimized multiplier tensor and the optimized control parameters are fed back to step S3 as the next moment state, and step S3 is repeated to update the gain control instruction.

[0049] In this embodiment, the augmented Lagrangian function is specifically formulated as follows:

[0050] Where, is the augmented Lagrangian function, is the objective function of the Lth layer, is the Lagrange multiplier of the Lth layer, is the constraint matrix of the Lth layer, is the control variable of the Lth layer, is the boundary vector of the Lth layer, is the preset penalty parameter.

[0051] In this embodiment, the objective function of the regional layer is to minimize the cross-layer power interaction cost, the objective function of the grid layer is to smooth load fluctuations, and the objective function of the unit layer is to minimize the load supply and demand deviation.

[0052] S5, according to the optimized gain control instructions, controls the execution units at each layer and coordinates the power flow, energy storage and load resources.

[0053] In this embodiment, the control of each layer execution unit coordinates power flow, energy storage and load resources, specifically: Controlling cross-layer power transmission capacity based on optimized gain control instructions; And based on the preset energy storage dynamic model, the energy storage charging power and energy storage discharging power of the grid layer are controlled; And control the adjustable load power of the unit layer.

[0054] In this embodiment, the specific formula of the preset energy storage dynamic model is:

[0055]

[0056] Where, is the energy storage capacity at time t, is the energy storage capacity at the moment before time t, and are the charging efficiency and discharging efficiency, and are the charging power and discharging power, and is the lower limit and upper limit of energy storage capacity, is the time interval.

[0057] It should be noted that cross-layer power transmission capacity refers to the maximum power limit allowed to be transmitted between different levels of power grids (such as regional layer, grid layer, and unit layer). It is used to ensure that the power interaction between layers does not exceed the carrying capacity of equipment or lines, while maintaining system stability. Furthermore, the reserve factor refers to the ratio of the reserve capacity reserved to cope with grid emergencies (such as equipment failure, sudden increase in load or sudden drop in renewable energy output) to the maximum load of the system; Furthermore, the self-sufficiency parameter refers to the proportion of the unit layer that can cover its maximum load through local distributed renewable energy and energy storage systems, which reflects the ability of local resources to support load demand.

[0058] The quantum long short-term memory (QLSTM) network is a neural network architecture that combines quantum computing with classical long short-term memory networks. Based on the classical LSTM, it replaces the neural network portion with variational quantum circuits, creating a classical-quantum neural network framework. By incorporating quantum computing properties such as coherent superposition and entanglement, the QLSTM can handle large amounts of data and complex pattern recognition tasks, enhancing feature extraction and data compression capabilities.

[0059] Singular spectrum analysis (SSA) is a signal processing and data analysis method based on time series, primarily used to extract trend terms, periodic terms, and noise components from time series. Its core concept is to construct a trajectory matrix and perform singular value decomposition (SVA) to decompose the original sequence into multiple subsequences with distinct characteristics, with each component corresponding to a distinct signal pattern or noise. SSA is model-independent and data-driven, making it suitable for tasks such as analysis, noise reduction, trend extraction, period identification, and forecasting of non-stationary time series.

[0060] The actor network is a core component of the reinforcement learning framework and a key part of the policy gradient method. Its primary function is to directly output a specific behavioral policy (i.e., a probability distribution of actions or deterministic actions) based on the environment state, optimizing its own policy through continuous interaction with the environment. The actor network works in conjunction with the critic network (value assessment network): the critic network evaluates the value of the actor's actions and provides feedback signals, which the actor then adjusts its policy parameters based on to maximize long-term cumulative rewards.

[0061] Reinforcement learning is a machine learning method that learns decision-making strategies through the interaction of an intelligent agent with its environment. The agent performs actions in the environment and adjusts its strategy based on the reward signals (positive or negative feedback) it receives to maximize long-term cumulative rewards. Its core elements include the agent, state, action, and reward, with key algorithms such as Q-learning, policy gradient, and actor-critic.

[0062] The alternating direction method of multipliers (ADMM) collaborative mechanism is an algorithmic framework for solving distributed optimization problems. The core idea is to decompose the complex global optimization problem into multiple sub-problems that can be solved in parallel by introducing auxiliary variables, and to achieve collaborative optimization between sub-problems using the Lagrange multiplier method and alternating update strategy. During the collaborative process, each sub-problem is optimized locally and independently, while being linked to other sub-problems through consistency constraints. The multiplier term is responsible for transmitting information between sub-problems and coordinating updates to ensure that local solutions gradually converge to the global optimum. This mechanism combines the high efficiency of distributed computing with the accuracy of centralized optimization. It is suitable for scenarios such as multi-agent systems, sensor networks, machine learning, and energy system optimization. In particular, it exhibits good scalability and convergence when dealing with large-scale, high-dimensional, and non-smooth objective functions, and can effectively balance computational complexity and collaborative efficiency.

[0063] The augmented Lagrangian function is an important tool for solving constrained optimization problems. By introducing a penalty term (or augmentation term) into the traditional Lagrangian function, the original problem is transformed into an unconstrained optimization problem, enhancing the algorithm's stability and convergence. Its basic form is to add a quadratic penalty term to the Lagrangian function, which is calculated based on the degree of constraint violation. This penalty term significantly impacts the objective function when the constraints are not satisfied, forcing the optimization process to more strictly satisfy the constraints. The augmented Lagrangian function is widely used in convex, non-convex, and distributed optimization algorithms. It is particularly effective in dealing with equality and inequality constraints, effectively balancing constraint satisfaction with objective function optimization, providing a flexible and efficient solution to complex optimization problems.

[0064] Example 2, Figure 2 The present invention provides an energy and power hierarchical probability balance control system based on a distribution grid, comprising: The scenario set construction module is used to predict the output of new energy sources based on the quantum-enhanced prediction algorithm and build a confidence scenario set in combination with the pre-acquired distribution model; The control parameter acquisition module is used to establish a three-level control architecture of region-grid-cell and obtain control parameters based on a set of confidence scenarios and preset hierarchical probability constraints; The gain control instruction acquisition module is used to generate gain control instructions based on control parameters and grid status data, with energy and power balance as the goal, through a policy-value collaborative reinforcement learning method; A gain control instruction update module is used to perform several rounds of iterative optimization on the gain control instruction based on a distributed collaborative optimization mechanism; The hierarchical control module is used to control the execution units at each layer according to the optimized gain control instructions and coordinate power flow, energy storage and load resources.

[0065] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0066] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0067] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0068] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0069] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0070] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for controlling energy and power hierarchical probability balance based on a distribution grid, characterized in that: The following steps are involved: Predicting new energy output based on quantum-enhanced prediction algorithms, combined with pre-acquired distribution models, to build a confidence scenario set; A three-level control architecture of region, grid, and cell is established to obtain control parameters based on a set of confidence scenarios and preset hierarchical probability constraints. Based on control parameters and grid status data, with energy and power balance as the goal, gain control instructions are generated through a policy-value collaborative reinforcement learning method; Perform several rounds of iterative optimization on the gain control instructions based on the distributed collaborative optimization mechanism; According to the optimized gain control instructions, the execution units at each layer are controlled to coordinate power flow, energy storage and load resources.

2. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 1 is characterized in that: The control parameters include cross-layer power transmission capacity, backup coefficient and self-sufficiency parameter; the grid status data includes energy residual, power residual and energy storage data; the gain control instructions include energy gain coefficient and power gain coefficient of each layer.

3. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 2 is characterized in that: The method for constructing the confidence scene set is specifically as follows: Based on SSA, the pre-obtained output time series is decomposed and grouped according to singular values ​​to obtain several groups of principal component subsequences. The principal component subsequences are then converted back into time series to obtain the corrected output time series. The corrected output time series is used as training data to train a pre-acquired QLSTM output prediction model and perform predictions to obtain an output prediction sequence, where the output includes photovoltaic output and wind power output; Based on the output forecast sequence, and the assumption that photovoltaic output follows the Beta distribution and wind power output follows the Weibull distribution, a set of renewable energy output probability scenarios is obtained through Monte Carlo sampling. The 1-norm constraint and ∞-norm constraint are imposed on the probability scenario set, and the typical scenarios are screened through the K-means clustering algorithm to output the new energy output confidence scenario set.

4. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 3 is characterized in that: The method for obtaining the control parameters is specifically as follows: Based on the confidence scenario set and regional layer constraints, the cross-layer power transfer capacity is solved with the goal of internal power balance in each region. Preset backup coefficients and determine grid layer constraints; Obtain the distributed new energy capacity and energy storage capacity at the unit level, and determine the self-sufficiency parameters through pre-set confidence intervals based on the self-sufficiency constraints of the unit level.

5. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 4 is characterized in that: The method for obtaining the gain control instruction is specifically as follows: Mapping control parameters and grid state data into a standardized state vector through a state encoder; According to the standardized state vector, based on the PPO algorithm, the energy gain coefficient and power gain coefficient of each layer are output through the Actor network, and the gain coefficient is mapped to the [-1, 1] interval through the tanh function; According to the mapped gain coefficient, the action value is calculated through the Critic network, and combined with the preset reward function, the Actor and Critic networks are updated through the policy gradient to output the gain control instruction.

6. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 5 is characterized in that: The gain control instruction is optimized for several rounds of iterations, specifically: Based on the ADMM collaborative mechanism, the augmented Lagrangian function is constructed in the digital twin environment, the control parameters are used as control variables and the multipliers are initialized; According to the energy residual and power residual, the control variables and multipliers of each layer are updated alternately, and the maximum number of iterations is preset until the residual converges or the maximum number of iterations is reached, and the tensor of the optimized multipliers and the optimized control parameters are obtained; The grid state data, the optimized multiplier tensor and the optimized control parameters are used as the state at the next moment, and the gain control instruction is updated.

7. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 6 is characterized in that: The control of each layer execution unit coordinates power flow, energy storage and load resources, specifically: Controlling cross-layer power transmission capacity based on optimized gain control instructions; And based on the preset energy storage dynamic model, the energy storage charging power and energy storage discharging power of the grid layer are controlled; And control the adjustable load power of the unit layer.

8. The energy and power hierarchical probability balance control method based on the distribution grid according to claim 7 is characterized in that: The updating of the Actor and Critic networks is specifically as follows: Update the Actor network through backpropagation with the goal of minimizing energy and power imbalance; The Critic network is updated with the goal of minimizing the temporal difference error of the Q value.

9. A system using the distribution grid-based energy and power hierarchical probabilistic balance control method according to any one of claims 1 to 8, comprising: The scenario set construction module is used to predict the output of new energy sources based on the quantum-enhanced prediction algorithm and build a confidence scenario set in combination with the pre-acquired distribution model; The control parameter acquisition module is used to establish a three-level control architecture of region-grid-cell and obtain control parameters based on a set of confidence scenarios and preset hierarchical probability constraints; The gain control instruction acquisition module is used to generate gain control instructions based on control parameters and grid status data, with energy and power balance as the goal, through a policy-value collaborative reinforcement learning method; A gain control instruction update module is used to perform several rounds of iterative optimization on the gain control instruction based on a distributed collaborative optimization mechanism; The hierarchical control module is used to control the execution units at each layer according to the optimized gain control instructions and coordinate power flow, energy storage and load resources.

Citation Information

Patent Citations

  • A dynamic energy efficiency optimization control method for MGT-CCHP under variable operating conditions

    CN114371619B

  • Distributed multi-agent system optimal control method based on strategy gradient

    CN116699979A