A power distribution network hierarchical optimization method for source-network-load interaction
By constructing a dynamic adaptive collaborative architecture and a hybrid intelligent solution strategy, the problems of insufficient adaptive capability and imbalance of interests in the distribution network under dynamic changes and multi-stakeholder interaction are solved, and efficient and stable distribution network optimization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FOSHAN GUYUXUAN BRAND MANAGEMENT CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-06-02
AI Technical Summary
The existing hierarchical architecture of the power distribution network lacks adaptability, suffers from imbalance in interest coordination, and is disconnected from planning and operation when facing dynamic changes and interactions among multiple stakeholders. This results in poor system robustness and low resource utilization efficiency under extreme disturbances.
A dynamic adaptive collaborative architecture is constructed, employing a three-layer hybrid game mechanism and a hybrid intelligent solution strategy. Combining multi-agent deep reinforcement learning and an improved particle swarm optimization algorithm, it achieves real-time hierarchical reconstruction, multi-interest coordination, and full-cycle optimization.
It enhances the system's adaptive reconfiguration capability under extreme disturbances, achieves precise coordination and fair allocation of resources among multiple stakeholders, breaks down decision-making barriers between planning and operation, and improves computational efficiency and control stability.
Smart Images

Figure CN121566432B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system automation technology, and more specifically, to a hierarchical optimization method for distribution networks oriented towards source-grid-load interaction. Background Technology
[0002] With the deepening of energy structure transformation and the construction of new power systems, distribution networks are rapidly evolving from traditional unidirectional passive networks into multidirectional interactive energy hubs integrating distributed power sources, flexible loads, energy storage devices, and intelligent control units. Against this backdrop, "source-grid-load" interaction has become a core path to improve the flexibility, resilience, and economy of distribution networks. To address the complexity brought about by high-proportion renewable energy integration, increased load fluctuations, and the participation of multiple stakeholders, academia and engineering generally adopt hierarchical coordination control and optimization strategies. The basic idea is to decompose the global high-dimensional nonlinear optimization problem into several clearly defined sub-problems with distinct responsibilities through spatiotemporal decoupling and functional decomposition, thereby reducing computational complexity while enhancing the system's responsiveness and robustness. Typical solutions are often based on a three-tiered "cloud-edge-device" architecture, where the cloud completes global optimization decisions, edge nodes perform regional coordination, and terminal devices are responsible for local responses. This paradigm has effectively improved scheduling efficiency and resource utilization in specific operating scenarios.
[0003] However, with the increasingly dynamic operating environment of distribution networks, the increasingly diversified interests of stakeholders, and the growing need for full-cycle coordination of "planning-operation-control," the aforementioned traditional layered architecture reveals deep-seated structural contradictions at the principle level. Specifically, its fixed hierarchical boundaries and static role allocation mechanisms essentially stem from neglecting the system's dynamic evolution capabilities. When the distribution network encounters topology reconfiguration, operating mode switching (such as grid-to-islanding), or extreme disturbance events, the pre-set control authority transfer logic is unable to adapt to the rapidly changing physical state and information flow pattern, leading to a lack of local autonomy or command conflicts, thereby weakening overall collaborative efficiency. Furthermore, the existing architecture lacks refined modeling of the complex electrical coupling relationships and information interaction logic between "source-grid-load-storage," often simplifying the coordination between multiple agents into one-way command transmission, ignoring the dynamic migration requirements of autonomous decision-making rights of each unit under different operating conditions, resulting in a lack of endogenous elastic reconfiguration capabilities when facing uncertain shocks. Ultimately, this defect does not stem from insufficient algorithm precision or communication bandwidth limitations, but rather from the imbalance in the architecture design itself in handling the inherent tension between "physical system dynamism" and "information control flexibility"—that is, in pursuing structural clarity, the system's adaptive reconfiguration capability under complex disturbances is sacrificed.
[0004] Meanwhile, the game relationships among diverse stakeholders exhibit highly mixed and nested characteristics, far exceeding the capabilities of existing single game models. Distribution network operators, microgrid clusters, distributed resource owners, and load aggregators are intertwined with both master-slave relationships and cooperative sharing intertwined with market competition, forming a dynamically evolving, multi-level game ecosystem. Using only a purely cooperative or purely competitive game framework not only fails to accurately reflect the actual interaction mechanisms but also easily leads to optimization results being difficult to implement due to imbalances in benefit distribution. Furthermore, current research generally treats planning, operation, and control as separate, independent processes, resulting in a disconnect between long-term investment decisions (such as energy storage deployment and line expansion) and short-term operational scheduling. The planned physical assets may fail to realize their expected regulatory potential in actual operation due to a lack of flexibility, leading to resource misallocation and economic losses. At the solution level, when faced with hierarchical optimization models that are high-dimensional, non-convex, and highly uncertain, traditional metaheuristic algorithms are prone to getting stuck in local optima and converge slowly. While emerging deep reinforcement learning methods have the advantage of sequential decision-making, they face real bottlenecks such as difficulty in embedding physical constraints, scarcity of training samples, and insufficient security verification, making it difficult to achieve efficient solutions while ensuring system security.
[0005] Therefore, how to construct a hierarchical optimization architecture that can dynamically adapt to the evolution of the distribution network's operating status, accurately depict the mixed game relationship among multiple stakeholders, and connect the entire cycle of "assessment-planning-operation," while integrating a hybrid intelligent solution mechanism with both strategy learning capabilities and strong constraint processing capabilities, in order to achieve a synergistic improvement in security, economy, and flexibility in high-penetration source-grid-load interaction scenarios, has become a key challenge and an urgent technical problem for those skilled in the art. Summary of the Invention
[0006] This invention provides a hierarchical optimization method for distribution networks oriented towards source-grid-load interaction, in order to solve the structural contradictions in the existing technology, such as insufficient system adaptability, imbalance of interests and inefficient resource utilization caused by rigid hierarchy, single game model and fragmented planning and operation.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0008] A hierarchical optimization method for distribution networks oriented towards source-grid-load interaction includes:
[0009] A dynamic adaptive collaborative architecture is constructed, which models the controllable units in the distribution network as intelligent agents with perception, decision-making and execution capabilities. Based on the real-time operating status of the distribution network, topological connection relationship and communication link quality, a three-layer logical structure of global coordination layer, regional autonomy layer and local execution layer is dynamically generated.
[0010] Based on the aforementioned dynamic adaptive collaborative architecture, a three-layer hybrid game mechanism is implemented, which includes a master-slave game between the main distribution network operator and the microgrid alliance, a cooperative game within the microgrid alliance, and a bilateral interactive game between distributed users and the microgrid, in order to determine the optimal interaction strategy for each entity.
[0011] The three-level closed-loop optimization process of "evaluation-planning-operation" is implemented. At the planning level, an initial plan is generated with the goal of minimizing the annual comprehensive cost. This plan is then passed to the operation level for multi-scenario rolling scheduling simulation. The key performance indicators fed back from the operation level are used as the basis for sensitivity analysis and fed back to the planning level to revise the plan until the preset comprehensive index system of operational flexibility is met.
[0012] A hybrid intelligent solution strategy is adopted to solve the three-level closed-loop optimization process. Multi-agent deep reinforcement learning is used to handle high-level coordination decision-making and target boundary generation. An improved particle swarm optimization algorithm is used to solve discrete optimization subproblems within a given boundary. Online interaction and parameter updates between the two are realized through a standardized interface.
[0013] Furthermore, the construction of the dynamic adaptive collaborative architecture specifically includes:
[0014] The configuration agent's state perception module continuously collects local electrical quantity information, and the role determination engine determines the logical level to which the agent belongs based on a preset threshold;
[0015] When the system experiences topology reconfiguration, island switching, or major disturbance events, a role redistribution mechanism for agents within the affected area is triggered.
[0016] The role redistribution mechanism adopts an election strategy based on multi-dimensional comprehensive scoring: if the microgrid switches to islanded operation mode due to a main tie-line failure, the intelligent agents with master control potential in the region calculate a comprehensive score based on remaining capacity, communication link quality, and computing resource margin and broadcast it. The one with the highest score is automatically promoted to the new regional autonomous layer master control intelligent agent, takes over the coordination and scheduling authority of other intelligent agents in the region, and reports the status change information of the island boundary, available adjustment resources list, and initial operating point to the global coordination layer.
[0017] Furthermore, the method of the present invention also includes establishing an event-triggered information interaction mechanism and a robust cooperative controller:
[0018] The event-triggered information interaction mechanism is configured with a hysteresis dead zone. When the local deviation detection unit detects that the key operating parameters deviate from the preset threshold range and the duration exceeds the anti-jitter time limit, or when it receives an external event signal, the information interaction process is initiated.
[0019] A cyber-physical fusion model considering communication non-ideals is constructed, jointly modeling the power network topology, equipment dynamic characteristics, and communication network latency, packet loss rate, and bandwidth limitations. A robust cooperative controller is designed, with the following control law:
[0020]
[0021] in, To control the input, The physical system state estimated by the state observer. This is the disturbance term estimated in real time by the communication error observer. and These are the state feedback gain and the disturbance compensation gain, respectively.
[0022] Furthermore, the master-slave game in the three-layer hybrid game mechanism is specifically as follows:
[0023] The main distribution network operator is the leader, setting the marginal electricity price and ancillary service price for each node; the microgrid alliance is the follower, optimizing its internal generation plan, energy storage charging and discharging power and controllable load reduction based on price signals, and feeding back the aggregated net power demand to the main distribution network operator.
[0024] Both parties solve the Stackelberg equilibrium iteratively to minimize the sum of total power purchase cost and network loss cost of the main distribution network while satisfying the power flow equations and security constraints, and at the same time maximize the net revenue of the microgrid alliance within the feasible region.
[0025] Furthermore, the cooperative game within the microgrid alliance in the three-layer hybrid game mechanism specifically refers to:
[0026] Alliance members are dynamically networked based on geographical proximity, resource complementarity, and operational reliability indicators. The total alliance revenue is distributed using an improved Shapley value method, with the following distribution formula:
[0027]
[0028] in, For members Distributed profits, The collection of all members within the microgrid consortium. S represents any specific member in N; S represents a subset of N that does not contain any members. Any subset of the microgrid consortium; Represents a set Any member index in; This is a return adjustment pool for redistribution after the introduction of a correction factor; For the characteristic function of the alliance; and This is a risk-sharing factor, and its value is positively correlated with the corresponding member's reserve capacity ratio. and The adjustment capability weight is positively correlated with the maximum adjustment rate of the corresponding member.
[0029] Furthermore, the bilateral interactive game in the three-layer hybrid game mechanism is specifically as follows:
[0030] Users participate in demand-side management by signing load response contracts. Their response behavior is jointly constrained by the price elasticity coefficient and comfort constraints, satisfying the following relationship:
[0031]
[0032] and ,in, and users respectively During the period The baseline load and the actual response load, This is the price elasticity coefficient. The incentive price issued for microgrids The benchmark electricity price, The maximum load adjustment acceptable to the user; the microgrid dynamically adjusts according to the user's response characteristics. To form the optimal incentive strategy under Nash equilibrium.
[0033] Furthermore, the aforementioned three-level closed-loop optimization process of "assessment-planning-operation" specifically includes:
[0034] A comprehensive index system for operational flexibility is defined at the evaluation level, including net load adaptability, regulation capacity margin, voltage support capability index, and line load balance.
[0035] At the planning level, the objective function is to minimize the annual comprehensive cost. The decision variables include the new line capacity, the location and capacity of energy storage configuration, and the access point and capacity of distributed power sources. Flexibility constraints are embedded in the model, requiring the scheme to meet the threshold requirements of the comprehensive index system of operational flexibility under typical operating scenarios.
[0036] The annual comprehensive cost includes new investment costs, operation and maintenance costs, and expected network loss costs.
[0037] Furthermore, the closed-loop optimization process also includes feedback corrections to the planning scheme from the runtime layer:
[0038] The runtime layer constructs a set of typical runtime scenarios based on the K-means clustering method, and performs multi-timescale rolling optimization, including day-ahead, intraday, and real-time, in each scenario;
[0039] The key operational results from the operation layer are fed back to the planning layer. These key operational results include the maximum power deficit, line congestion frequency, and voltage over-limit count.
[0040] Based on the feedback results, the planning layer uses sensitivity analysis to identify weak links and executes a correction process until the planning scheme meets the flexibility requirements in all typical scenarios and the annual comprehensive cost change rate is less than the preset value. Each execution of the correction process includes: if the frequency of line congestion in a target area is continuously higher than the threshold, priority is given to expanding the tie line in the target area; if the number of voltage over-limit times is concentrated at the end node, then the configuration of distributed energy storage is increased.
[0041] Furthermore, the multi-agent deep reinforcement learning specifically adopts the MADRL framework:
[0042] Each autonomous layer agent is configured with an independent deep Q-network, with the input being the state observation vectors of the local and neighboring regions, and the output being the scheduling command setpoint;
[0043] The reward function is designed as a composite function, which includes the cost of electricity purchase, the cost of grid loss, and penalties for voltage and power overruns.
[0044] The training process uses a centralized experience replay pool to store the transition samples of all agents and a distributed policy update mechanism, where each agent independently calculates the gradient but shares some network parameters.
[0045] Furthermore, the improved particle swarm optimization algorithm and its interaction with multi-agent deep reinforcement learning include:
[0046] For subproblems containing a large number of discrete variables, an improved particle swarm optimization algorithm with a multi-subgroup cooperative mechanism is adopted to solve the problem. The population is divided into multiple subgroups for independent search, and elite individuals are migrated after each iteration. Dynamically decaying inertial weights are introduced, and simulated annealing acceptance criteria are integrated to allow the acceptance of inferior solutions with a specific probability in the later stages of iteration.
[0047] The interaction process is conducted through a standardized interface: the multi-agent deep reinforcement learning outputs a boundary update instruction every preset time interval, and the improved particle swarm optimization algorithm reinitializes the population and performs local optimization accordingly; if the improved particle swarm optimization algorithm cannot find a feasible solution within the given boundary, it triggers the boundary relaxation mechanism and feeds back the degree of constraint violation as a penalty signal, forcing the multi-agent deep reinforcement learning to adjust the policy search space in the next time step; if a feasible solution is found, the optimal objective function value is fed back for policy evaluation.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] (1) It breaks through the rigid constraints of the traditional hierarchical architecture and improves the system's adaptive reconfiguration capability to extreme disturbances. This invention constructs a dynamic adaptive collaborative architecture, which no longer divides the hierarchy based on fixed physical devices, but dynamically generates logical hierarchy based on the real-time operating status, topology, and communication quality of the distribution network. When topology reconfiguration or island switching occurs, in conjunction with the role redistribution mechanism based on multi-dimensional comprehensive scoring, it can achieve a smooth transfer of control authority, solve the problem of lack of local autonomy or instruction conflict in the face of sudden situations in the existing fixed hierarchy, and significantly enhance the robustness of the system.
[0050] (2) It achieves precise coordination and fair allocation of complex interests among multiple stakeholders. The three-layer hybrid game mechanism proposed in this invention covers the master-slave game between the main grid and microgrids, the cooperative game between microgrids, and the bilateral interaction on the user side. In particular, the introduction of the improved Shapley value method, which comprehensively considers the risk-sharing factor and the adjustment capacity weight for the distribution of benefits, reflects both the marginal contribution and the heterogeneity of the members. It effectively solves the problem that the scheduling strategy is difficult to implement due to the difficulty of balancing the interests of all parties in a single game model, and incentivizes high-quality resources to actively participate in grid interaction.
[0051] (3) It breaks down the decision-making barriers between planning and operation, establishing a full-cycle closed-loop optimization model. Through the three-level closed-loop optimization framework of "evaluation-planning-operation", this invention abandons the traditional one-way process that separates planning and operation. The operation layer feeds back key indicators such as power deficit and congestion frequency in actual scheduling to the planning layer, using sensitivity analysis to guide the precise expansion or energy storage configuration of weak links. This mechanism ensures that the planning scheme has sufficient flexibility in actual operation, avoids asset idleness or insufficient adjustment capacity due to neglecting operational details, and achieves the dual optimization of investment benefits and operational safety.
[0052] (4) This invention integrates the decision-making advantages of deep reinforcement learning with the constraint handling capabilities of metaheuristic algorithms, achieving efficient solutions in complex scenarios. For high-dimensional, nonlinear, and strongly coupled optimization problems, this invention employs a hybrid intelligent solution strategy. Multi-agent deep reinforcement learning (MADRL) is used to handle the strong uncertainty generated by high-level policies, while an improved particle swarm optimization (PSO) algorithm is used to solve the optimization of discrete variables under specific constraints. Furthermore, boundary relaxation and penalty feedback mechanisms are used to achieve synergy between the two. This not only overcomes the shortcomings of traditional algorithms that are prone to getting trapped in local optima, but also solves the problem that purely data-driven methods cannot strictly satisfy physical constraints, significantly improving computational efficiency and convergence stability.
[0053] (5) Enhanced system control stability under non-ideal communication environments. This invention explicitly considers the impact of communication delay and packet loss on physical states by establishing a cyber-physical fusion model and designing a robust cooperative controller. Combined with an event triggering mechanism with dead time hysteresis, it reduces communication load while ensuring accurate execution of control commands in communication-constrained scenarios, thus guaranteeing the stability of distribution network voltage and frequency.
[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, embodiments of the present invention are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a schematic diagram of the overall architecture of the distribution network hierarchical optimization method oriented towards source-grid-load interaction as described in this invention.
[0057] Figure 2 This is a schematic diagram of the three-layer logical structure of the dynamic adaptive collaborative architecture and the intelligent agent role switching mechanism in this invention.
[0058] Figure 3 This is a schematic diagram of the interaction relationship and information flow of the three-layer hybrid game mechanism adopted in this invention.
[0059] Figure 4 This is a flowchart illustrating the "evaluation-planning-operation" three-level closed-loop optimization framework of this invention.
[0060] Figure 5 This is a schematic diagram illustrating the collaborative working mechanism of deep reinforcement learning and improved particle swarm optimization algorithm in the hybrid intelligent solution strategy of this invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0062] This invention provides a hierarchical optimization method for distribution networks oriented towards source-grid-load interaction. Its core lies in constructing a collaborative optimization system with dynamic adaptive capabilities, a multi-stakeholder interest coordination mechanism, and integrated planning and operation characteristics.
[0063] First, the overall system architecture is as follows: Figure 1 As shown, this invention consists of four core modules: a dynamic adaptive collaborative architecture, a three-layer hybrid game mechanism, an "evaluation-planning-operation" closed-loop optimization framework, and a hybrid intelligent solution strategy. These modules are tightly coupled through standardized data interfaces and event-driven mechanisms, forming a scalable, reconfigurable, and highly robust distribution network collaborative optimization platform.
[0064] Specifically, the hierarchical optimization method for distribution networks oriented towards source-grid-load interaction described in this invention mainly includes the following steps:
[0065] Step 1: Construct a dynamic adaptive collaborative architecture, model the controllable units in the distribution network as intelligent agents with perception, decision-making and execution capabilities, and dynamically generate a three-layer logical structure of global coordination layer, regional autonomy layer and local execution layer based on the real-time operating status of the distribution network, topological connection relationship and communication link quality.
[0066] Step 2: Based on the dynamic adaptive collaborative architecture, a three-layer hybrid game mechanism is run to determine the optimal interaction strategy of each subject through master-slave game between the main distribution network operator and the microgrid alliance, cooperative game within the microgrid alliance, and bilateral interactive game between distributed users and the microgrid.
[0067] Step 3: Execute the three-level closed-loop optimization process of "evaluation-planning-operation". At the planning level, an initial plan is generated with the goal of minimizing the annual comprehensive cost. The plan is then passed to the operation level for multi-scenario rolling scheduling simulation. The key performance indicators fed back from the operation level are used as the basis for sensitivity analysis and fed back to the planning level to revise the plan.
[0068] Step four: A hybrid intelligent solution strategy is adopted to solve the three-level closed-loop optimization process. Multi-agent deep reinforcement learning is used to handle high-level coordination decision-making and target boundary generation, and an improved particle swarm optimization algorithm is used to solve discrete optimization sub-problems within the given boundary.
[0069] The dynamic adaptive collaborative architecture adopts a three-layer logical structure, namely the global coordination layer, the regional autonomy layer, and the local execution layer, such as... Figure 2 As shown, this three-layer structure is not based on fixed physical equipment or geographical location, but is dynamically generated according to the real-time operating status of the distribution network, topological connections, communication link quality, and the control capabilities of each unit. All controllable units within the system, including but not limited to distributed photovoltaic inverters, energy storage converters, flexible load aggregation controllers, and microgrid central controllers, are modeled as intelligent agents with sensing, decision-making, and execution capabilities. Each intelligent agent integrates three functional sub-modules: a state perception module, a role determination engine, and an instruction execution unit.
[0070] The state awareness module continuously collects local electrical quantity information, including voltage amplitude, current phase, active and reactive power, frequency deviation, and net load fluctuation rate, and simultaneously receives external inputs such as superior dispatch instructions, market electricity price signals, and protection action events. The role determination engine has a built-in rule set that determines the logical level to which the current intelligent agent should belong based on preset thresholds and logical conditions. In normal grid-connected operation mode, distributed photovoltaic inverters and ordinary load controllers typically belong to the local execution layer; energy storage converters and flexible load aggregation controllers, due to their bidirectional regulation capabilities, can be assigned to the regional autonomous layer; while the microgrid central controller defaults to being the master intelligent agent of the regional autonomous layer, responsible for coordinating the operating strategies of all execution layer intelligent agents within the region.
[0071] When the system experiences topology reconfiguration, islanding switching, or major disturbances, agents within the affected area trigger a role redistribution mechanism through local state monitoring. This mechanism employs a multi-dimensional comprehensive scoring-based election strategy: if a microgrid is forced into islanding operation due to a main tie-line fault, agents within the area with master control potential (such as energy storage converters) will calculate a comprehensive score based on their remaining capacity (SOC), communication link quality, and computing resource margin, and broadcast it. The agent with the highest score is automatically promoted to the new regional autonomous layer master control agent; if there are identical scores, a unique master control agent is determined based on a preset device ID priority or a random backoff algorithm, taking over the coordination and scheduling authority of other agents within the area, and reporting state change information, including island boundaries, available adjustable resources, and initial operating point, to the global coordination layer via the uplink communication channel. Upon receiving this information, the global coordination layer updates the entire network logical topology model and adjusts the coordination strategy for adjacent areas to avoid cross-regional power backfeeding or voltage instability.
[0072] To reduce communication load and improve response efficiency, this invention employs an event-triggered information interaction mechanism with a dead time hysteresis. Each agent is equipped with a local deviation detection unit to continuously monitor key operating parameters. When any parameter deviates from a preset threshold range and the duration exceeds the anti-jitter time limit (e.g., node voltage amplitude exceeds ±5% of the nominal value, branch power exceeds 90% of the thermal stability limit, frequency deviation is greater than ±0.2Hz, net load fluctuation rate exceeds twice the standard deviation of the historical average), or when an external event signal (such as a relay protection action signal) is received, the agent initiates the information interaction process.
[0073] The direction of information interaction is dynamically determined based on the event type and the agent's role. For example, when a local execution layer agent detects a voltage limit violation, it sends a data packet containing a timestamp, state vector (voltage, power, SOC, etc.), constraint boundaries (maximum charge / discharge power, regulation rate), and objective function gradient information to the higher-level regional autonomous layer master agent. After receiving multiple lower-level requests, the regional autonomous layer master agent performs local optimization calculations and uploads the aggregated net power demand and regulation capability boundaries to the global coordination layer. Simultaneously, regional autonomous layers at the same level can also exchange boundary information to coordinate power flow in the tie-line. All data packets are encapsulated in a unified format to ensure that the receiver can quickly parse and participate in collaborative optimization calculations, avoiding interaction delays caused by protocol incompatibility.
[0074] Furthermore, to quantify the impact of communication non-ideals on control performance, this invention simultaneously constructs a cyber-physical fusion model. This model jointly models the power network topology, equipment dynamic characteristics, and communication network latency, packet loss rate, and bandwidth limitations. Considering the nonlinear characteristics of power flow in the distribution network, small-signal linearization is performed near the system's steady-state operating point, defining t as a continuous-time variable; let the power system's incremental state vector be... ,in express A real vector space; let the state vector of the communication system be... ,in express In a real vector space, the coupled state space equation can be expressed as:
[0075]
[0076] in, The physical system state matrix, For the state matrix of the communication system, and These are the coupling matrices between communication and physics, and between physics and communication, respectively. and These are the input matrices for the physical system and the communication system, respectively. and The coupling effect of information flow on the physical system and the physical state on communication scheduling is characterized. Based on this, a robust cooperative controller is designed, with the following control law:
[0077]
[0078] in, The physical system state estimated by the state observer. This is the disturbance term estimated in real time by the communication error observer. and These are the state feedback gain and disturbance compensation gain, respectively. The controller is deployed within the regional autonomous layer's master control agent to ensure that hierarchical optimization instructions are accurately executed even in the event of data loss or transmission delay, maintaining system voltage, frequency, and power balance within safe limits.
[0079] Regarding the coordination of interests among multiple stakeholders, this invention constructs a three-layer hybrid game mechanism, such as... Figure 3 As shown, the top layer is a master-slave game structure, with the backbone distribution network operator as the leader, setting the marginal electricity price for each node. Price of ancillary services The microgrid consortium, as a follower, optimizes its internal [mechanism / system] based on price signals. Power generation plan for the time period Energy storage charging and discharging power and controllable load reduction and the calculated first Net power demand during the period Feedback is then sent to the main distribution network operator. Through iterative solutions to the Stackelberg equilibrium, both parties ensure that the main distribution network minimizes the sum of total electricity purchase cost and network loss cost while satisfying power flow equations, line capacity constraints, and voltage safety constraints, while the microgrid consortium maximizes its net revenue within its feasible region.
[0080] The second layer is the cooperative game mechanism within the microgrid alliance. Alliance members dynamically form networks based on geographical proximity, resource complementarity, and operational reliability indicators. The total alliance revenue... The allocation is determined by each member's contribution and uses a modified Shapley value method. This method introduces a risk-sharing factor and a moderating capacity weight to the traditional Shapley value for secondary correction, ensuring that the allocation result reflects both marginal contribution and the heterogeneity of the members. The specific allocation formula is as follows:
[0081]
[0082] in, The collection of all members within the microgrid consortium. Represents a set Any member index in; This is a revenue adjustment pool for redistribution after the introduction of the correction factor (which can be set as a certain percentage of the total alliance revenue). This is the characteristic function of the alliance. and It is a risk-sharing factor (its value is positively correlated with the reserve capacity ratio of the corresponding member). and The adjustment capacity weights (their values are positively correlated with the maximum adjustment rate of the corresponding member). The denominator term in the formula... This is used to normalize the weighted product of all members. This allocation mechanism ensures that members with greater adjustment potential or who bear higher operational risks receive a higher share of the profits, thereby incentivizing high-quality resources to actively participate in the alliance's operation.
[0083] The third layer is the bilateral interaction mechanism between distributed users and the microgrid. Users participate in demand-side management by signing load response contracts, and their response behavior is influenced by the price elasticity coefficient. It is constrained in conjunction with comfort constraints. Let the user... During the period The baseline load is The actual response load is Then the following condition is met:
[0084]
[0085] and ,in The incentive price issued for microgrids The benchmark electricity price, This represents the maximum load adjustment acceptable to the user. The microgrid dynamically adjusts based on user response characteristics. This leads to the formation of the optimal incentive strategy under Nash equilibrium. The three-layer game mechanism is integrated through a unified mathematical model to form a mixed-integer nonlinear programming problem. Its solution adopts the improved alternating direction multiplier method (ADMM), in which each party only needs to exchange Lagrange multipliers and coupling variables (such as tie-line power and nodal price), without disclosing private cost functions or load curves, effectively protecting business privacy.
[0086] Regarding the coordination of planning and operation, this invention constructs a three-level closed-loop optimization framework of "assessment-planning-operation," such as... Figure 4 As shown. The evaluation layer first defines a comprehensive index system for operational flexibility, including net load adaptability. Adjusting capacity margin Voltage support capability index and line load balance Among them, the net load adaptability rate is defined as the probability that the system can balance net load fluctuations without shedding loads under typical disturbance scenarios; the regulation capacity margin is the ratio of available upward / downward adjustment capacity to the maximum prediction deviation; the voltage support capability index is the reciprocal of the root mean square value of voltage deviation at all nodes; and the line load balance is the reciprocal of the standard deviation of the load rate of each branch. This indicator system serves as a common optimization objective for both the planning and operation levels, ensuring consistent guidance for decisions at all levels.
[0087] The planning layer uses the minimum annual comprehensive cost as its objective function, expressed as:
[0088]
[0089] in, For the planning period, For the discount rate, For the first Annual increase in investment costs (including line expansion, energy storage configuration, and distributed power source access). For maintenance costs, The expected network loss cost. Decision variables include the capacity of new lines. Energy storage configuration location With capacity Distributed power access point and capacity The planning model incorporates flexibility constraints, requiring the selected solution to meet certain conditions under typical operating scenarios. , Equal threshold, where and These are the preset lower limit thresholds for net load adaptability and adjustment capacity margin, respectively.
[0090] After the initial planning scheme is generated, it is passed to the operation layer for multi-scenario scheduling simulation. Based on historical meteorological data, load curves, and equipment fault records, the operation layer uses K-means clustering to construct a set of typical operation scenarios, covering all four seasons (spring / summer / autumn / winter), three weather types (sunny / cloudy / rainy), and two load patterns (weekday / weekend), generating a total of 24 typical scenarios. Under each scenario, the operation layer performs multi-timescale rolling optimization: In the day-ahead phase, a basic scheduling plan is generated based on 72-hour net load and renewable energy output forecasts, with optimization variables including unit start-up and shutdown, energy storage charging and discharging plans, and tie-line power; every 4 hours during the intraday phase, the day-ahead plan is revised based on the rolling 24-hour forecast, adjusting the flexible load response; and every 5 minutes during the real-time phase, rapid correction control is performed based on SCADA measurement data, primarily adjusting the energy storage converter output power and the inverter reactive power setpoint.
[0091] Each timescale optimization aims to maximize flexibility or minimize operating costs, and incorporates key operational results (such as maximum power deficit) into the optimization process. Frequency of line congestion Number of times voltage exceeds limit Feedback is sent to the planning layer. Based on the performance indicators fed back from the operation layer, the planning layer uses sensitivity analysis to identify weak points, such as if a certain area... If the value remains above the threshold, priority will be given to expanding the interconnection lines in that area; if If energy storage is concentrated at the end nodes, distributed energy storage configuration is increased. This iterative process continues until the planning scheme meets the flexibility requirements in all typical scenarios and the annual comprehensive cost change rate is less than 0.1%, or the number of iterations reaches a preset upper limit (e.g., 20 times). If the iteration upper limit is reached but the scheme has not fully converged, the scheme with the best comprehensive evaluation index in each iteration is selected as the final result. The final scheme obtained at this time combines economy and operational adaptability, avoiding the problems of asset idleness or insufficient capacity caused by neglecting operational details in traditional planning.
[0092] Regarding the solution strategy, this invention employs a hybrid solution architecture combining deep reinforcement learning and improved metaheuristic algorithms, such as... Figure 5 As shown, high-level coordination decisions are accomplished using a multi-agent deep reinforcement learning (MADRL) framework. Each regional autonomous layer agent is configured with an independent deep Q-network (DQN), with the input being the state observation vectors of its local and neighboring regions. Where the subscript m represents the autonomous agent of the m-th region, This represents the voltage amplitude of the node where the intelligent agent is located. Active power Reactive power In the state of energy storage charge, For the marginal electricity price at the node, For frequency deviation; output is the scheduling command setpoint. ,in This refers to the adjustment of active power for energy storage. This is the reactive power adjustment value of the inverter. This represents the load shedding amount. The reward function is designed as a composite form:
[0093]
[0094]
[0095] in, The unit price of electricity. Costs are calculated based on network loss. Let be the reward value obtained by the m-th agent. The amount of electricity purchased by agent m from the upper-level power grid. For network losses in this area, This represents the actual current flowing through the critical lines in this area. This is the rated current value of the line; the last two terms of the formula are penalty terms for voltage and power overruns (or line overloads), respectively. and These are the penalty coefficients for exceeding voltage and power limits, respectively. The training process uses a centralized experience replay pool to store the transfer samples of all agents. ,in This is the current state observation vector (i.e., the state observation vector mentioned earlier). The action to be executed at the current moment (i.e., the scheduling instruction setpoint mentioned earlier). This is the reward value obtained previously. The state observation vector is generated after the action is performed, and a distributed policy update mechanism is adopted. Each agent independently calculates the gradient but shares some network parameters to improve sample utilization efficiency and convergence stability.
[0096] For subproblems involving a large number of discrete variables, such as energy storage start-up and shutdown decisions Network switch operation For equipment selection (such as energy storage type selection), an improved particle swarm optimization (PSO) algorithm is used. This algorithm introduces a multi-subgroup cooperative mechanism, dividing the population into... Each subgroup independently searches different regions of the solution space; after each iteration, the best individuals in each subgroup are sorted by fitness, with the top... Elite individuals migrate to other subgroups to achieve information sharing. The inertia weight employs a dynamic decay strategy.
[0097]
[0098] in, Let the current iteration algebra be... The maximum number of iterations, For the first The inertia weight value at the next iteration and These represent the maximum and minimum values of the inertia weight, respectively. Simulated annealing acceptance criteria are also incorporated, allowing for probabilistic acceptance in later iterations. Accept inferior solutions, among which The annealing temperature. This represents the initial temperature of the annealing algorithm. The temperature decay coefficient is , The degree of fitness degradation is represented by the algorithm. This algorithm is deployed in each local execution layer agent to quickly solve mixed-integer nonlinear subproblems within their jurisdiction.
[0099] In the hybrid solution architecture, the deep reinforcement learning agent is responsible for generating high-level coordination policies and objective boundaries (such as upper and lower limits of energy storage SOC and tie-line power constraints), while the metaheuristic solver is responsible for solving specific discrete optimization problems within the given boundaries. The two interact online through a standardized interface: MADRL outputs a boundary update instruction every 15 minutes, which the improved PSO algorithm uses to reinitialize the population and perform local optimization. If PSO cannot find a feasible solution within the given boundaries, a boundary relaxation mechanism is triggered, and the degree of constraint violation is fed back to MADRL as a penalty signal, forcing it to adjust its policy search space in the next time step. If a feasible solution is found, the optimal objective function value is fed back to MADRL for policy evaluation in the next cycle. This mechanism forms a closed-loop learning process of 'policy generation-execution verification-feedback correction', combining the long-term sequential decision-making advantages of deep reinforcement learning with the direct handling capability of metaheuristic algorithms for complex constraints.
[0100] This invention achieves real-time reconstruction of the logical hierarchy through a dynamic adaptive collaborative architecture, accurately coordinates the interests of multiple stakeholders through a three-layer hybrid game mechanism, connects the entire lifecycle decision-making chain through an "evaluation-planning-operation" closed-loop framework, and balances solution efficiency and engineering feasibility through a hybrid intelligent solution strategy. These technical elements support and organically integrate each other, forming a hierarchical optimization method for distribution networks in scenarios with high proportions of source-grid-load interaction, providing a complete technical implementation path for the efficient, safe, and flexible operation of new power systems.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A hierarchical optimization method for distribution networks oriented towards source-grid-load interaction, characterized in that, include: A dynamic adaptive collaborative architecture is constructed, which models the controllable units in the distribution network as intelligent agents with perception, decision-making and execution capabilities. Based on the real-time operating status of the distribution network, topological connection relationship and communication link quality, a three-layer logical structure of global coordination layer, regional autonomy layer and local execution layer is dynamically generated. Based on the aforementioned dynamic adaptive collaborative architecture, a three-layer hybrid game mechanism is implemented, which includes a master-slave game between the main distribution network operator and the microgrid alliance, a cooperative game within the microgrid alliance, and a bilateral interactive game between distributed users and the microgrid, in order to determine the optimal interaction strategy for each entity. The three-level closed-loop optimization process of "evaluation-planning-operation" is implemented. At the planning level, an initial plan is generated with the goal of minimizing the annual comprehensive cost. This plan is then passed to the operation level for multi-scenario rolling scheduling simulation. The key performance indicators fed back from the operation level are used as the basis for sensitivity analysis and fed back to the planning level to revise the plan until the preset comprehensive index system of operational flexibility is met. A hybrid intelligent solution strategy is adopted to solve the three-level closed-loop optimization process. Multi-agent deep reinforcement learning is used to handle high-level coordination decision-making and target boundary generation. An improved particle swarm optimization algorithm is used to solve discrete optimization sub-problems within a given boundary. Online interaction and parameter updates between the two are realized through a standardized interface. The multi-agent deep reinforcement learning specifically adopts the MADRL framework: Each autonomous layer agent is configured with an independent deep Q-network, with the input being the state observation vectors of the local and neighboring regions, and the output being the scheduling command setpoint; The reward function is designed as a composite function, which includes the cost of electricity purchase, the cost of grid loss, and penalties for voltage and power overruns. The training process uses a centralized experience replay pool to store the transition samples of all agents and a distributed policy update mechanism, in which each agent independently calculates the gradient but shares some network parameters. The improved particle swarm optimization algorithm and its interaction with multi-agent deep reinforcement learning include: For subproblems containing discrete variables, an improved particle swarm optimization algorithm with a multi-subgroup cooperative mechanism is used to solve the problem. The population is divided into multiple subgroups for independent search, and elite individuals are migrated after each iteration. Dynamically decaying inertial weights are introduced, and simulated annealing acceptance criteria are incorporated to allow the acceptance of inferior solutions with a specific probability in the later stages of iteration. The interaction process is conducted through a standardized interface: the multi-agent deep reinforcement learning outputs a boundary update instruction every preset time interval, and the improved particle swarm optimization algorithm reinitializes the population and performs local optimization accordingly; if the improved particle swarm optimization algorithm cannot find a feasible solution within the given boundary, it triggers the boundary relaxation mechanism and feeds back the degree of constraint violation as a penalty signal, forcing the multi-agent deep reinforcement learning to adjust the policy search space in the next time step; if a feasible solution is found, the optimal objective function value is fed back for policy evaluation.
2. The hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 1, characterized in that, The construction of the dynamic adaptive collaborative architecture specifically includes: The configuration agent's state perception module continuously collects local electrical quantity information, and the role determination engine determines the logical level to which the agent belongs based on a preset threshold; When the system experiences topology reconfiguration, island switching, or major disturbance events, a role redistribution mechanism for agents within the affected area is triggered. The role redistribution mechanism adopts an election strategy based on multi-dimensional comprehensive scoring: if the microgrid switches to islanded operation mode due to a main tie-line failure, the intelligent agents with master control potential in the region calculate a comprehensive score based on remaining capacity, communication link quality, and computing resource margin and broadcast it. The one with the highest score is automatically promoted to the new regional autonomous layer master control intelligent agent, takes over the coordination and scheduling authority of other intelligent agents in the region, and reports the status change information of the island boundary, available adjustment resources list, and initial operating point to the global coordination layer.
3. The hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 2, characterized in that, This also includes establishing an event-triggered information interaction mechanism and a robust collaborative controller: The event-triggered information interaction mechanism is configured with a hysteresis dead zone. When the local deviation detection unit detects that the key operating parameters deviate from the preset threshold range and the duration exceeds the anti-jitter time limit, or when it receives an external event signal, the information interaction process is initiated. A cyber-physical fusion model considering communication non-ideals is constructed, jointly modeling the power network topology, equipment dynamic characteristics, and communication network latency, packet loss rate, and bandwidth limitations. A robust cooperative controller is designed, with the following control law: Where t is a continuous-time variable. To control the input, The physical system state estimated by the state observer. This is the disturbance term estimated in real time by the communication error observer. and These are the state feedback gain and the disturbance compensation gain, respectively.
4. The hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 1, characterized in that, The master-slave game in the three-layer hybrid game mechanism is specifically as follows: The main distribution network operator is the leader, setting the marginal electricity price and ancillary service price for each node; the microgrid alliance is the follower, optimizing its internal generation plan, energy storage charging and discharging power and controllable load reduction based on price signals, and feeding back the aggregated net power demand to the main distribution network operator. Both parties solve the Stackelberg equilibrium iteratively to minimize the sum of total power purchase cost and network loss cost of the main distribution network while satisfying the power flow equations and security constraints, and at the same time maximize the net revenue of the microgrid alliance within the feasible region.
5. A hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 1, characterized in that, The cooperative game within the microgrid alliance in the three-layer hybrid game mechanism is specifically as follows: Alliance members are dynamically networked based on geographical proximity, resource complementarity, and operational reliability indicators. The total alliance revenue is distributed using an improved Shapley value method, with the following distribution formula: in, For members Distributed profits, The collection of all members within the microgrid consortium. S represents any specific member in N; S represents a subset of N that does not contain any members. Any subset of the microgrid consortium; Represents a set Any member index in; This is a return adjustment pool for redistribution after the introduction of a correction factor; For the characteristic function of the alliance; and This is a risk-sharing factor, and its value is positively correlated with the corresponding member's reserve capacity ratio. and The adjustment capability weight is positively correlated with the maximum adjustment rate of the corresponding member.
6. The hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 1, characterized in that, The bilateral interactive game in the three-layer hybrid game mechanism is specifically as follows: Users participate in demand-side management by signing load response contracts. Their response behavior is jointly constrained by the price elasticity coefficient and comfort constraints, satisfying the following relationship: and ,in, and users respectively During the period The baseline load and the actual response load, This is the price elasticity coefficient. The incentive price issued for microgrids The benchmark electricity price, The maximum load adjustment acceptable to the user; the microgrid dynamically adjusts according to the user's response characteristics. To form the optimal incentive strategy under Nash equilibrium.
7. A hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 1, characterized in that, The three-level closed-loop optimization process of "assessment-planning-operation" specifically includes: A comprehensive index system for operational flexibility is defined at the evaluation level, including net load adaptability, regulation capacity margin, voltage support capability index, and line load balance. At the planning level, the objective function is to minimize the annual comprehensive cost. The decision variables include the new line capacity, the location and capacity of energy storage configuration, and the access point and capacity of distributed power sources. Flexibility constraints are embedded in the model, requiring the scheme to meet the threshold requirements of the comprehensive index system of operational flexibility under typical operating scenarios. The annual comprehensive cost includes new investment costs, operation and maintenance costs, and expected network loss costs.
8. A hierarchical optimization method for distribution networks oriented towards source-grid-load interaction as described in claim 7, characterized in that, The closed-loop optimization process also includes feedback and correction of the planning scheme by the runtime layer: The runtime layer constructs a set of typical runtime scenarios based on the K-means clustering method, and performs multi-timescale rolling optimization, including day-ahead, intraday, and real-time, in each scenario; The key operational results from the operation layer are fed back to the planning layer. These key operational results include the maximum power deficit, line congestion frequency, and voltage over-limit count. Based on the feedback results, the planning layer uses sensitivity analysis to identify weak links and executes a correction process until the planning scheme meets the flexibility requirements in all typical scenarios and the annual comprehensive cost change rate is less than the preset value. Each execution correction process includes: if the line congestion frequency in a target area continues to be higher than the threshold, priority is given to expanding the tie line capacity in that target area; If the number of voltage overruns is concentrated at the end nodes, then add distributed energy storage configuration.