Distributed resource optimization scheduling method for flexibility improvement

By constructing a dynamic incentive model and a deep autoencoder, and combining it with distribution network topology constraints, a two-layer master-slave game model is established, which solves the problems of excessive equipment wear and the curse of computational dimensionality in distributed resource scheduling, and achieves highly reliable power grid scheduling.

CN122178452APending Publication Date: 2026-06-09ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER
Filing Date
2026-05-11
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing distributed resource scheduling schemes are prone to problems such as excessive equipment wear and tear, the curse of computational dimensionality, local power flow exceeding limits, and the disconnect between the economic interests of multiple stakeholders and physical security when dealing with massive and heterogeneous resources. Traditional centralized scheduling is difficult to achieve highly reliable closed-loop scheduling equilibrium.

Method used

By constructing a dynamic excitation-based secondary decay mapping model, extracting latent features using a deep autoencoder, dividing virtual machine groups by combining distribution network topology constraints, establishing a multi-time-scale hierarchical scheduling architecture, and building a two-layer master-slave game model between the upper and lower layers to optimize resource scheduling.

Benefits of technology

It enables precise quantification of the power response boundary, reduces the computational dimension, ensures the executability of dispatch instructions and the physical security of the power grid, and improves the overall capacity for resource absorption efficiency and economic optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122178452A_ABST
    Figure CN122178452A_ABST
Patent Text Reader

Abstract

The application discloses a kind of distributed resource optimization scheduling methods for flexibility promotion, it is related to electric power system intelligent dispatching technical field.The method includes: obtaining heterogeneous resource multidimensional feature and regulation intensity, constructs dynamic incentive function and establishes secondary attenuation mapping model to quantify physical electric energy response boundary;Using deep auto-encoder to extract latent layer features, combine distribution network topology constraints to cluster resources into multiple virtual groups;Based on virtual group, establish multi-time scale scheduling architecture, and construct double-layer master-slave game model, the electric energy response boundary is used as lower physical constraint condition;Using the optimization algorithm of inertia weight dynamic adjustment introduced to solve the game model, output optimal dynamic incentive strategy and scheduling execution instruction.The present application is used to solve the problem that traditional linear scheduling is easy to cause equipment overconsumption, dimension disaster and local flow overrun caused by massive dispersed resources directly accessing, and the disconnection between multi-agent economic benefit pursuit and bottom physical safety bottom line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent dispatching technology for power systems, and more specifically, to a distributed resource optimization dispatching method aimed at improving flexibility. Background Technology

[0002] The high proportion of wind and solar power generation and complex loads integrated into new distribution networks has dramatically increased the randomness of system power fluctuations. Massive, heterogeneous distributed resources, with their potential for bidirectional power regulation, have become a core source of flexibility for ensuring the dynamic balance of the power grid. Efficiently aggregating and scheduling these massive, dispersed flexible resources is a pressing technical challenge in the field of power regulation.

[0003] Existing distributed resource scheduling schemes mostly adopt centralized control architecture and linear economic models. In terms of response mechanism, they usually rely on a fixed linear price sensitivity function to evaluate the adjustment potential of various resources; when aggregating and partitioning resources, they mainly rely on static clustering based on a single geographical or administrative boundary; at the decision optimization level, single-level programming models are generally used to pursue the minimization of global operating costs and directly solve for the output control commands.

[0004] However, existing technologies have significant limitations: First, linear sensitivity models are prone to severely overestimating resource regulation capabilities under extreme excitations, neglecting nonlinear physical boundaries such as equipment lifespan depletion and mechanical fatigue, resulting in scheduling instructions lacking practical executability. Second, facing massive influx of high-dimensional operational data, traditional centralized static scheduling is highly susceptible to the "curse of dimensionality" and real-time computational bottlenecks. It fails to effectively extract the latent complementarity of heterogeneous features and deviates from distribution network topology constraints, easily inducing local power flow exceedance risks. Finally, traditional single-layer planning struggles to coordinate the overall benefits of the upper layer with the conflicts of interest among multiple stakeholders at the lower layer, resulting in insufficient resource utilization. Furthermore, the economic optimization boundary is severely disconnected from the strong physical constraints at the lower level, making it difficult to achieve highly reliable closed-loop scheduling equilibrium. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a distributed resource optimization scheduling method for improving flexibility. Before game modeling, a dynamic incentive-based quadratic decay mapping model is constructed to quantify the real physical response boundary, a virtual machine group is aggregated using a deep autoencoder and topological constraints, and a multi-timescale, two-layer master-slave game model with strong constraints on the underlying power response boundary is established for dynamic optimization. This addresses the problems that traditional linear scheduling easily leads to excessive equipment wear, the curse of computational dimensionality caused by the access of massive distributed resources, and local power flow exceeding limits, as well as the disconnect between the pursuit of economic interests by multiple stakeholders and the bottom line of underlying physical security.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A distributed resource optimization scheduling method for improving flexibility includes the following steps: obtaining multi-dimensional feature vectors and adjustment intensity variables of heterogeneous distributed resources, and constructing a dynamic excitation function; constructing a quadratic attenuation mapping model based on the dynamic excitation function to quantify the power response boundary; extracting latent features of the multi-dimensional feature vectors through a deep autoencoder, and dividing the heterogeneous distributed resources into multiple virtual machine groups in combination with distribution network topology constraints; establishing a multi-time-scale hierarchical scheduling architecture based on the virtual machine groups; establishing a two-layer master-slave game model between the upper-layer scheduling unit and the lower-layer distribution system in the hierarchical scheduling architecture, with the power response boundary serving as the lower-layer physical constraint condition; solving the game model using an optimization algorithm that introduces dynamic adjustment of inertial weights, and outputting the optimal dynamic excitation strategy and scheduling execution instructions.

[0007] The technical effects and advantages of this invention's distributed resource optimization scheduling method for improving flexibility are as follows: Before game theory modeling, this invention first constructs a dynamic excitation function based on the adjustment intensity variable, then quantifies the power response boundary through a quadratic decay mapping model, and uses it as a physical constraint for lower-level resource scheduling, thus ensuring that the game solution naturally falls within the safe operating range of the equipment. This improvement solves the problem at the physical level that existing game theory models have superior economic objectives but unsafe actual scheduling instructions. Simultaneously, before game theory modeling, this invention extracts the latent complementarity of multidimensional heterogeneous resource features through a deep autoencoder, and combines this with the actual topological constraints of the distribution network to divide massive resources into several virtual machine groups, ensuring that the aggregated units reduce the optimization dimensionality without violating physical safety constraints such as distribution network line capacity. This is fundamentally different from existing resource aggregation strategies based on clustering or administrative boundary division—the aggregation process no longer deviates from the actual operation of the power grid, thus ensuring the executability of the game theory decision. Attached Figure Description

[0008] Figure 1 A schematic diagram of a distributed resource optimization scheduling method for improving flexibility provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of three-dimensional fitting of the power response provided in an embodiment of the present invention; Figure 3 A flowchart of the improved particle swarm optimization algorithm provided in an embodiment of the present invention; Figure 4 This is a comparison chart of the convergence performance of the algorithms provided in the embodiments of the present invention. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0010] Example 1, Figure 1 This invention presents a distributed resource optimization scheduling method for improving flexibility, comprising the following steps: Step S1: Obtain the multidimensional feature vectors and adjustment intensity variables of heterogeneous distributed resources, and construct a dynamic activation function.

[0011] In this embodiment, the step of obtaining the multidimensional feature vector and adjustment intensity variable of heterogeneous distributed resources and constructing a dynamic activation function specifically includes the following sub-steps: S101, Collect the rated power, operating status and excitation response sensitivity of the heterogeneous distributed resources, and combine them to form the multidimensional feature vector; Specifically, in actual power distribution network operation scenarios, the rated power and operating status of various heterogeneous distributed resources within the jurisdiction are simultaneously acquired through Wide Area Measurement System (WAMS), Supervisory Control and Data Acquisition (SCADA) systems, and IoT smart meters widely deployed at the terminal side. Specifically, the rated power of each resource is collected as a static baseline feature; simultaneously, the dynamic operating status of different resources is extracted. For energy-type equipment such as battery storage, its operating status includes real-time state of charge (SOC), current available charge / discharge capacity, and battery health parameters; for highly stochastic units such as wind power generation, its operating status includes real-time wind speed, generator speed, and the error degree of the current output deviation from the day-ahead forecast curve; for industrial controllable loads, it includes the load baseline of the current production process, minimum interruptible time, and recovery delay.

[0012] The formula for calculating the excitation response sensitivity is: (1) in, For the first The incentive response sensitivity of a distributed resource is the mathematical expectation of the absolute ratio of the change in actual response power to the change in price. The total number of historical periods to be evaluated. For the first The actual response power of each historical evaluation period For baseline power, and The first and The incentive prices for each historical period. Power and price data for these historical periods were obtained by extracting historical scheduling logs to measure the actual response rate and execution deviation of each resource at different price tiers in past scheduling. It should be noted that, during the calculation, incentive prices were excluded. To prevent computational overflow, invalid evaluation cycles are eliminated.

[0013] After acquiring the rated power, operating status, and excitation response sensitivity data, the data is cleaned and normalized using a deviation standardization method, mapping all feature parameters across dimensions to [0,1]. Subsequently, the data is concatenated according to a preset resource feature topology dimension to form a multidimensional feature vector. This serves as the standard quantitative input for subsequent deep-layer feature extraction.

[0014] The preset resource feature topology dimensions are concatenated, as shown in the following example: Assuming the system's preset standard feature dimension is 10, the first and second dimensions of the vector are fixed as common physical attributes shared by all resources (normalized rated power, excitation response sensitivity, etc.). The vector is structured as follows: dimensions 3 to 5 are fixed as energy storage-specific attributes (SOC, available capacity, health); dimensions 6 to 8 are fixed as wind power-specific attributes (real-time wind speed, generator speed, prediction error); and dimensions 9 to 10 are fixed as load-specific attributes (baseline load, minimum interruptible time). When processing data for a specific resource (such as a battery energy storage system), its corresponding feature quantization values ​​are filled into dimensions 1 to 5 of the vector. For wind power and load-specific attribute positions (i.e., dimensions 6 to 10) that the energy storage system does not possess, zero values ​​are used as placeholders.

[0015] Further, step S102 is performed to obtain the adjustment intensity variable. ,in ; As a preferred embodiment, the adjustment intensity variable It is a dynamic physical quantity characterizing the urgency of demand for regulation in a macro-level power grid under a specific dispatch section. Its value is generated in real-time by the upper-level power grid dispatch center based on the current system's safety margin, net load peak-to-valley difference rate, rate of frequency sag (RoCoF), or the degree of congestion in the local distribution network. Specifically, it can be calculated by normalizing various multi-source over-limit assessment indicators and then performing a weighted summation. Its mathematical model is: (2) in, To evaluate the total number of dimensions of the indicators, For the first Real-time quantitative values ​​of various evaluation metrics (such as normalized peak-to-valley ratio or congestion level). Let be the dynamic weight coefficient of the corresponding indicator, and satisfy . .

[0016] when When the value approaches 0, it indicates that the current system is in a normal operating range with ample supply and demand and stable power flow, mainly adopting a low-intensity voluntary adjustment strategy based on information guidance. As the system faces extreme boundary conditions such as sudden drops in wind and solar power output or surges in load, the pressure for peak shaving or emergency frequency regulation increases. When the value approaches 1, it indicates that the need for adjustment is extremely urgent, and a high-intensity intervention mechanism with strong rigid constraints and a high call price must be triggered.

[0017] After quantifying the distribution network regulation demand, step S103 is executed to construct the dynamic excitation function: (3) The specific physical meanings of each variable are as follows: This represents the output dynamic incentive price signal, which will serve as the core driving parameter for regulating the response behavior of the lower-level distributed units. This indicates the maximum incentive price ceiling that the system can withstand within the current scheduling cycle, subject to the overall economic budget and market rules. The rate of increase coefficient characterizes the sensitivity of incentive funding to changes in urgency; a higher value indicates a more dramatic price surge when demand crosses a critical point. In a preferred embodiment, the rate of increase coefficient... The range of values ​​is set to This is to ensure that when the demand for regulation crosses the inflection point, the price can be quickly stimulated and reach the threshold state. The inflection point parameter of the function curve represents the threshold state at which the incentive strategy shifts from moderate economic guidance to strong price-driven stimulus.

[0018] In a preferred embodiment, the inflection point parameter The range of values ​​is set to This ensures that the system maintains high economic efficiency under normal fluctuations, and only triggers high incentives under more urgent boundary conditions.

[0019] The dynamic excitation function accurately simulates the "saturation effect" in economics and control: adjusting the intensity... In lower, soft-guided phases or extremely high, mandatory-directed phases, the slope of incentive price changes gradually flattens to ensure the effectiveness of incentive funds, avoid ineffective over-subsidization of inefficient responses, and strictly adhere to the economic boundaries of the distribution network; while... Approaching and crossing the inflection point During the critical ramp-up phase, incentive prices rise exponentially, thereby maximizing the response potential of massive distributed resources within the budget.

[0020] This step addresses the technical shortcomings of traditional linear excitation models, which struggle to accurately reflect the capability boundaries of underlying equipment under extreme conditions, leading to a lack of physical executability of control commands and low system control precision. By implementing step S1, a quantitative characterization of the regulation capability of massive heterogeneous distributed resources is achieved, resolving the problem of unclear regulation characteristics under dispersed resource conditions. This method employs a Logistic function with nonlinear characteristics to construct a dynamic excitation model and combines it with multidimensional resource characteristics to build a nonlinear "generalized excitation-adjustable capability" mapping model. This model successfully establishes a correlation mechanism between the macro-control demand of the power grid and microeconomic price signals, thus providing solid data support and model basis for subsequent accurate quantification of power response boundaries and the implementation of hierarchical cluster scheduling.

[0021] Step S2, constructing a quadratic decay mapping model based on the dynamic excitation function, and quantifying the electrical energy response boundary, including: The dynamic excitation function Using the output as the independent variable, quadratic function response models incorporating nonlinear decay characteristics and boundary capacity constraints are constructed for battery energy storage systems, wind power generation systems, and industrial controllable loads, respectively, to obtain the corresponding electrical energy response quantities. .

[0022] In actual power distribution network regulation, the response patterns of distributed resources under different price incentive signals exhibit significant differences and nonlinear marginal benefit changes. To accurately quantify the regulatory effect of price signals on flexible and diverse resources, this embodiment links the underlying physical lifetime and topological constraints of resources with economic responses, and constructs corresponding quadratic function response models of "electricity response-price incentive" for different resource types, including: battery energy storage system response model, wind power generation system response model, and industrial controllable load response model.

[0023] (1) Battery energy storage system response model For battery energy storage systems, the charge and discharge response is fast and relatively flexible. However, as incentive prices rise, frequent deep charge and discharge cycles will trigger the battery's internal state of charge (SOC) safety limit protection, leading to accelerated degradation of cell cycle life and usable capacity. Therefore, it is necessary to establish an energy storage response capacity that includes lifespan constraints and physical boundaries. Model: (4) In the formula, The dynamic excitation function The output incentive price signal; It is the response decay coefficient of the battery energy storage system under high price incentives. It is generally greater than that of pumped hydro storage and characterizes the battery's rapid response capability under high price (considering cycle life and SOC control). The linear sensitivity of battery energy storage characterizes the system's normal rapid response and following capability under low-to-medium price incentives, and is more sensitive to incentive prices compared to pumping. The minimum amount of energy to be drawn from the battery, such as the initial response under SOC limits; This represents the maximum physical response power boundary allowed by the current SOC and health status. When the incentive price... When the response exceeds its extreme point, the model uses upper and lower limiting mechanisms to maintain the electrical response within safe physical boundaries. The operation avoids unreasonable fitting caused by the parabolic descent interval.

[0024] (2) Response model of wind power generation system For wind power systems, although a certain reserve can be reserved through active power control (APC), excessive exploitation of peak-shaving response capacity under extreme high-wind conditions or in pursuit of high incentive returns can easily induce turbine pitch angle adjustment saturation and mechanical fatigue damage to the drivetrain. To avoid overloading and exceeding the lifespan of wind turbines under high-price incentives, its response power... The model is constructed as follows: (5) In the formula, The nonlinear attenuation coefficient is the factor that limits the lifespan of the wind turbine under excessive excitation due to equipment life protection or mechanical strength limitation. This represents the response incentive willingness coefficient of the wind power system. The available basic power supply is constrained by the real-time wind speed and the reference speed of the wind turbine in the current environment. This represents the maximum permissible output boundary under the current environmental real-time wind speed and equipment safety limits. Similarly, when the incentive price... When the response exceeds its extreme point, the model uses a limiting mechanism to forcibly truncate the electrical response and maintain it within safe physical boundaries. run.

[0025] (3) Industrial controllable load response model For controllable industrial loads, their response potential is constrained by the rigid limitations of plant thermodynamic processes or safe production procedures, resulting in limited response within this scale. However, under extremely high excitation, advance process scheduling may lead to a leap in the response quantity. Therefore, its response power... The model is constructed as follows: (6) In the formula, The nonlinear driving response coefficient of industrial load characterizes the special nonlinear increment brought about by process flow adjustment. Linear price sensitivity of industrial load; The physical boundary of the basic adjustable load that can be directly disconnected from the grid or participate in regulation during non-production periods; This represents the maximum interruptible load for the current production shift at the factory.

[0026] In this embodiment, the above , and The model coefficients can be obtained by extracting the scatter plot data of the actual response power of the corresponding equipment in different price ranges from the historical scheduling database and then using nonlinear least squares curve fitting.

[0027] This step corrects the bias in the traditional linear scheduling model's assessment of distributed resource regulation capabilities under extreme adjustment demands by constructing a quadratic function response model. This mapping mechanism transforms physical boundary conditions such as equipment lifespan depletion, state of charge (SOC) limitations, and mechanical fatigue into model constraints, reducing the risks of overcharging, over-discharging, and fatigue damage to distributed resources, thereby improving the physical safety and system reliability during scheduling instruction execution.

[0028] Step S3: Extract the latent features of the multidimensional feature vector through a deep autoencoder, and divide the heterogeneous distributed resources into multiple virtual machine groups in combination with the distribution network topology constraints.

[0029] In this embodiment, the extraction of latent features from the multidimensional feature vector using a deep autoencoder specifically includes the following sub-steps: S301, the multidimensional feature vector is input into the deep autoencoder, and the encoder maps the multidimensional feature vector to the latent space to obtain the latent layer features.

[0030] Specifically, since the heterogeneous resource multidimensional features collected in step S1 include highly coupled and redundant heterogeneous data such as rated power, operating status, and price sensitivity, directly using them as scheduling input would cause an exponential expansion of the computational scale. Therefore, this method constructs a deep autoencoder network containing an input layer, multiple fully connected hidden layers, and a BottleNeck layer (bottleneck layer) for nonlinear dimensionality reduction.

[0031] During forward propagation, the multidimensional feature vectors enter the input layer and are compressed and extracted layer by layer by the encoder. Finally, they are mapped to the BottleNeck layer with the fewest nodes, forming low-dimensional latent features that reflect the essential operating mode of the resources. The ReLU activation function is used for the neuron activation mechanism in the hidden layers of the network to avoid the gradient vanishing problem in deep network training and to enhance the model's ability to express strongly nonlinear power data.

[0032] To ensure that the compressed latent features do not lose key physical adjustment parameters, the autoencoder network is trained unsupervised with the goal of minimizing the reconstruction error. Its loss function is calculated as follows: (7) In the formula, This represents the total loss function value during the training process of the autoencoder network; The high-dimensional original multidimensional feature vector received by the input layer; This is the feature matrix output by the decoder after the feature vector is reconstructed. The reconstruction error term characterizes the degree of feature deviation between the input data and the reconstructed data; The reduced-dimensional latent layer representation features are output by the BottleNeck layer; This is a regularization term applied to the latent features to ensure that the latent feature matrix remains sparsity; The weighting coefficients are used to control the strength of regularization.

[0033] Deep autoencoders are trained in an unsupervised manner with the goal of minimizing reconstruction error. This ensures that the latent features output by the encoder can retain information related to the regulation capability to the greatest extent from the original multidimensional feature vector (which includes key physical regulation parameters such as rated power, operating status, and stimulus response sensitivity). This avoids the loss of physical features that are crucial for subsequent virtual machine group partitioning and response boundary quantization due to dimensionality reduction, thus providing an accurate and complete input basis for the generalized clustering distance metric built based on latent features.

[0034] To ensure the sparsity of latent features and highlight the core physical patterns, the regularization term is penalized using the L1 norm, which is specifically calculated as follows: ,in Output vectors for BottleNeck layer The Each dimension component This represents the total number of nodes in the BottleNeck layer.

[0035] Further, sub-step S302 is executed, whereby clustering is performed based on the latent features and combined with the distribution network topology constraints to complete the partitioning of virtual machine groups: After obtaining low-dimensional latent features characterizing the core attributes of distributed resources, an improved hierarchical clustering algorithm is used to dynamically aggregate massive, dispersed resources into several appropriately sized entity units. To ensure the safety constraints of the clustering results in the actual physical power grid, this embodiment abandons the single Euclidean distance when calculating the distance between different resource nodes in the latent space. Instead, it introduces the electrical impedance matrix and node voltage sensitivity matrix of the distribution network as quantitative penalty terms for topological correlation, constructing the following generalized clustering distance measurement model: (8) In the formula, For distributed resources With resources Generalized clustering distance between them; and For the corresponding low-dimensional latent layer features, The difference in data characteristics between the two in terms of operating modes such as response speed and adjustable capacity; Resources in the distribution network impedance matrix and The electrical impedance values ​​between the nodes represent the electrical span of the topology; and The voltage sensitivity matrix elements represent the nodes. Voltage on resources and Sensitivity parameters for power injection; and This represents the maximum adjustable output force boundary for both. This refers to the set of vulnerable nodes that are the focus of key monitoring in the power distribution network. and These are the impedance penalty weight and the voltage over-limit penalty weight, respectively, and their values ​​are determined by per-unit conversion of the distribution network reference voltage and reference impedance. It is an exponential penalty function. In a preferred embodiment, the sensitivity amplification factor is used. The range of values ​​is , For nodes The voltage safety limit threshold is defined. By introducing this exponential term, the severe penalty effect when approaching the safety boundary can be accurately reflected.

[0036] The specific physical logic constraint quantization operation mechanism is as follows: when the algorithm evaluates resources... and When merging virtual machines into the same virtual machine cluster, if the electrical distance between the two is too great (i.e., impedance), When the output is extremely high (spanning multiple transformer substations), or when both simultaneously respond to high-price incentives and achieve peak output, their combined effect can lead to a vulnerable node. If the expected voltage deviation value approaches the safety limit threshold, the corresponding exponential penalty term will be amplified exponentially, thus significantly increasing the generalized clustering distance between the two in the algorithm space. This forces the termination of the possibility of coordinated discharge between the two under the same scheduling command. Conversely, for resources with similar topological locations and complementary spatiotemporal response characteristics (e.g., energy storage units with second-level response paired with controllable loads with long-term regulation), the generalized clustering distance is minimized, thus they are tightly clustered, ensuring that resources within each cluster have similar regulation characteristics and complementary operating modes.

[0037] When performing hierarchical clustering, the uniform connectivity method is used to calculate the distance between clusters, and a truncation threshold is set. When the generalized clustering distance between two merged clusters is greater than 1, the merging distance is greater than 1. When the iteration aggregation stops, the resulting independent clusters are the partitioned virtual machine groups.

[0038] Traditional aggregation methods often employ Euclidean distance clustering based on response characteristics, neglecting the electrical correlation of resources within the physical network. This leads to the incorrect aggregation of geographically proximate resources on different feeders, easily inducing local voltage exceedances or transformer overloads during coordinated responses. This new method introduces a generalized clustering distance that incorporates penalties for electrical impedance and voltage sensitivity, embedding hard physical topology constraints into the clustering criteria in situ. This forcibly cuts off resource aggregation paths that may induce safety risks, achieving resource cluster partitioning that balances the complementarity of response modes with the security of the physical network. It solves the problem of local power flow exceedances caused by ignoring grid topology constraints during large-scale distributed resource aggregation.

[0039] By implementing step S3, this invention utilizes a dimensionality-reduced clustering mechanism based on electrical topology constraints and sensitivity quantization to reduce the communication load and computational complexity when massive discrete resources participate in centralized optimization scheduling. It achieves reduced-order equivalent representation of multi-source heterogeneous resources. By introducing physical network topology constraints during the aggregation process, it effectively reduces the risk of local feeder power flow exceeding limits and voltage deviation, ensuring the safety and feasibility of multi-timescale scheduling commands at the power grid physical level.

[0040] S4. Establish a multi-time-scale hierarchical scheduling architecture based on the virtual machine group.

[0041] In this embodiment, establishing a multi-time-scale hierarchical scheduling architecture based on the virtual machine group specifically includes the following sub-steps: S401, Day-ahead scheduling phase: Two-stage stochastic planning based on scenario analysis is adopted to formulate the basic power plan for each virtual machine group; Specifically, during the day-ahead dispatch phase with a 24-hour dispatch cycle and a time resolution of 15 minutes or 1 hour, the system faces the severe challenge of both fluctuations in renewable energy output and uncertainties in market electricity prices. This embodiment adopts a two-stage stochastic programming framework: the first stage is the pre-decision-making of day-ahead output and market bidding, requiring each virtual machine group to submit a baseline operating plan before the occurrence of uncertainty; the second stage is the real-time physical balancing stage. In order to effectively avoid the high penalties caused by large-scale wind and solar curtailment or the shedding of critical loads due to extreme wind and solar deviations, Conditional Value at Risk (CVaR) is introduced to accurately measure the tail risk of dispatch exceeding limits.

[0042] The two-stage objective function that considers risk preference is formulated as follows: (9) In the formula, This represents the optimal combined target value for expected operating costs and risk penalties in the current phase. These are the decision variables for the first stage, namely the pre-planned base power for each virtual machine group; Stochastic scenario parameters characterizing wind and solar power output and load prediction errors; To account for the expected scheduling cost considering the probability of each operating scenario, where The scheduling cost function for heterogeneous distributed resources; Indicates at confidence level The conditional value at risk is used to quantify the extreme value of expected loss in the worst-case scenario. This is the risk aversion coefficient for dispatchers regarding uncertainty. The higher the value, the stronger the redundancy of daytime output.

[0043] Among them, random scenarios Typical discrete scene sets of wind and solar power output were generated using Latin hypercube sampling (LHS) combined with K-means clustering. To facilitate the solution, auxiliary variables were introduced. and scene tail loss variables Non-smooth Equivalent to linear constraints: (10) In the formula, For scenario probabilities, This represents the total number of discrete scenarios. In a preferred embodiment, the confidence level is... The value is 0.95, representing the risk aversion coefficient. The value range is [0.5, 2.0].

[0044] Further, in sub-step S402, the intraday rolling scheduling phase is executed: model predictive control is adopted to track the energy interaction benchmark results while satisfying the linearized power flow constraints of the distribution network. During the intraday operation phase, to eliminate ultra-short-term forecast bias, a rolling control framework is activated with a period of 5-15 minutes. A model predictive control (MPC) algorithm is employed, by setting a longer forecast time domain. With a relatively short control time domain The quadratic programming problem is solved within progressively pushed-back time windows to achieve smooth tracking of the day-ahead base power plan.

[0045] To ensure the absolute physical safety of the underlying power grid when power commands are issued to virtual machine groups, a linearized DistFlow model is forcibly introduced as a hard power flow constraint when solving the MPC objective function, ensuring that node voltage and branch power do not exceed limits. The formulas for tracking the optimization objective and the core power flow node voltage constraints are as follows: (11) (12) In the formula, Based on the current time The predicted future period will yield results; The energy interaction reference power issued in step S401; Output increments for intraday control orders; and These are the weight matrices representing the tracking error penalty and the control action smoothness penalty, respectively. In the DistFlow constraint, and These represent the operating voltages at the beginning and end nodes of the distribution network branch, respectively. This is the system's reference rated voltage; and For nodes The feeder resistance and reactance between them; and This refers to the active and reactive power transmission through this branch. This constraint ensures that every intraday rolling tracking action is strictly locked within safe voltage and capacity boundaries. In intraday rolling optimization, the prediction time domain is set. (i.e., predicting 1 hour ahead), controlling the time domain Weight matrix It is a unit diagonal matrix to enhance the tracking of the reference power.

[0046] As a preferred implementation, sub-step S403 is executed for transient extreme operating conditions, in the real-time control stage: when a sudden fault or frequency voltage exceedance is detected in the distribution network, a high-sensitivity node is located using a sensitivity matrix, and the heterogeneous distributed resources corresponding to the high-sensitivity node are controlled to inject reactive power to provide voltage support.

[0047] Specifically, in the ultra-real-time control range of seconds to milliseconds, when the underlying measurements detect an excessively high system rate of change (RoCoF) or a deep voltage drop at critical nodes, conventional intraday control is no longer sufficient to suppress the deteriorating trend in time. At this point, by solving the inverse matrix of the current distribution network power flow Jacobian matrix, the voltage-reactive power sensitivity matrix between nodes is derived in real time. This precisely identifies the high-sensitivity source nodes that provide the most significant support to the over-limit nodes. Specifically, by sorting the coefficients of the target nodes in the sensitivity matrix in descending order, the top N nodes are selected as the execution nodes. Subsequently, an over-limit trigger is initiated at the selected high-sensitivity execution nodes, where the grid-connected inverter (GFM) performs an emergency autonomous response. Using the sensitivity matrix, a mapping relationship is established between the reactive power injection and the desired voltage support. The voltage support quantification formula is as follows: (13) In the formula, For nodes where sudden faults or voltage drops occur The voltage deviation to be achieved through reactive power support; These are the corresponding elements within the voltage sensitivity matrix constructed in step S3 above, representing the control node. For the faulty node The voltage affects the gain coefficient; For nodes The reactive power support amount urgently injected by heterogeneous distributed resources in a network structure with a response time of seconds.

[0048] Furthermore, the sensitivity formula described above clarifies the reactive power deficit that needs to be compensated. In real-time control, to achieve closed-loop self-healing, nodes... The GFM at the location adopts a droop control strategy based on local voltage deviation to adaptively inject reactive power.

[0049] In one optional implementation, the step of injecting reactive power into the heterogeneous distributed resources corresponding to the high-sensitivity node to provide voltage support specifically includes: using the heterogeneous distributed resources corresponding to the high-sensitivity node as a grid-type inverter, and adaptively injecting reactive power using a droop control law based on reactive capacity margin adaptive adjustment; the control law is specifically: (14) in, Based on the droop control coefficient For nodes The current actual reactive power generated by the grid-connected inverter. To constrain the physical boundary of the maximum available reactive power due to the equipment's rated parameters, As a margin penalty adjustment factor, For nodes Reference rated voltage, For real-time voltage measurement.

[0050] In conventional fixed droop factor logic, when faced with a deep voltage drop, each node will generate reactive power response to the same extent, which can easily cause inverters with small reactive power capacity margins (i.e. poor basic conditions) to become instantly overloaded and saturated and lose support.

[0051] This embodiment incorporates a nonlinear penalty term in series within the underlying closed-loop control. An adaptive assessment and allocation module based on capacity margin was constructed. Under actual operating conditions, when the inverter's actual reactive power... When the margin is small and sufficient, the penalty term approaches 0, and the inverter operates close to the base coefficient. The strong response slope injects reactive power; and as the inverter's output continues to increase and approaches the physical boundary... At that time, the penalty term rapidly amplifies non-linearly and approaches 1 (margin penalty adjustment factor). The optimal setting is between 2 and 4, which smoothly reduces the equivalent droop control coefficient to near 0.

[0052] Through the aforementioned adaptive dynamic adjustment mechanism based on underlying physical parameters, devices approaching their physical capacity limits automatically slow down their response slope. This allows the reactive power support pressure on the local power grid to be spontaneously and smoothly transferred to adjacent nodes with more sufficient margin through network-level electrical coupling. This mechanism effectively prevents secondary grid disconnection caused by overload of a single-point inverter at the underlying physical boundary of the electrical topology, significantly improving the high reliability and physical security of the local distribution network's autonomous voltage closed-loop self-healing mechanism.

[0053] This method constructs a multi-timescale hierarchical scheduling architecture of day-to-day and real-time through step S4. The virtual machine group maintains the same partitioning structure under different time scales, which provides a stable decision granularity and time connection basis for the subsequent two-layer master-slave game model and avoids coordination difficulties caused by frequent changes in grouping.

[0054] Step S5: In the hierarchical scheduling architecture, a two-layer master-slave game model is established between the upper-layer scheduling unit and the lower-layer power distribution system, and the power response boundary is used as the lower-layer physical constraint condition.

[0055] It should be noted that the two-layer master-slave game model constructed in step S5 is the core of the specific economic optimization and market clearing in the day-ahead scheduling pre-decision stage of the preceding step S401. Specifically, when formulating the basic power plan for each virtual machine group in the day-ahead stage, multiple uncertain and random scenarios are encountered. For each specific scenario, the system calls the two-layer master-slave game model described in this step (S5) to find the optimal solution for the interaction of electricity and price. The optimal interaction of electricity and operating cost fed back by the lower-level power distribution system in the game model will serve as the underlying basis for calculating the expected scheduling cost of the scenario in step S401; while the optimal dynamic incentive strategy output by the upper-level leader constitutes the basic execution plan for day-ahead scheduling.

[0056] In this embodiment, establishing a two-layer master-slave game model between the upper-layer scheduling unit and the lower-layer power distribution system specifically includes the following sub-steps: S501, Construct a dynamic incentive model for upper-level leaders, with the objective function being to maximize the scheduling revenue of the upper-level scheduling unit within a scheduling cycle. Under the constraints of inter-grid power anti-cyclic constraints and main grid-side benchmark incentive boundaries, formulate time-sharing energy interaction incentive signals to be issued to each of the lower-level distribution systems. Specifically, in the new distribution network architecture involving multiple stakeholders, the upper-level dispatching unit (such as the Distributed Resource Aggregator (DAR)) and the lower-level Flexible Distribution System (FDS) belong to different stakeholders. As the leader in the Stackelberg game, the upper-level dispatching unit aims to guide the response of lower-level stakeholders by optimizing its own published electricity purchase and sale prices, thereby maximizing net profit by leveraging the price differences between the upper and lower levels and its interaction with the main grid. Its profit-maximizing objective function is constructed as follows: (15) In the formula, For the upper-level scheduling unit in a complete scheduling cycle (e.g.) Total revenue within (hours); This represents the total number of lower-level power distribution systems (FDS). and The upper layer in time period To the The electricity sales price and purchase price set by each lower-level power distribution system (i.e., time-of-use energy interaction incentive signal); and The first The lower-level power distribution system purchases and sells electricity to the upper-level system; and The benchmark on-grid electricity price and the grid sales price are the main grid side (upper-level electricity market). and This refers to the electricity sold and purchased by the upper-level scheduling unit in transactions with the main network.

[0057] Meanwhile, the upper-level scheduling unit at any time period The power balance constraint at the aggregator level must be satisfied, and its energy conservation equation is: (16) This equation ensures that the total electricity purchased by the upper layer from the main grid and the lower layer is equal to the total electricity sold.

[0058] To prevent lower-level entities from engaging in malicious arbitrage by exploiting time differences in electricity prices or variations across multiple nodes, and to ensure the authenticity and effectiveness of energy flow, the model introduces inter-grid power anti-cycle constraints: (17) This formula mandates that only one-way power exchange can occur within the same lower-level node at the same time period, and that simultaneous power purchase and sale are strictly prohibited, thus eliminating the ineffective idling of power at the underlying logic level.

[0059] Meanwhile, to ensure the lower-level distribution system's participation in this layer's game without being disconnected from the grid for trading, the set incentive signals must satisfy the main grid's baseline incentive boundary constraints: (18) That is, the set electricity purchase and sale prices must be strictly limited to the range between the grid electricity price and the on-grid electricity price.

[0060] Further, sub-step S502 is executed to construct a lower-level follower collaborative optimization model. The objective function is to minimize the operating cost of each lower-level power distribution system. Under the conditions of satisfying node power balance constraints, load shift conservation constraints, and energy storage system charging and discharging mutual exclusion constraints, the output and interactive power of internal equipment are optimized.

[0061] The lower-level power distribution system, as a follower in the game, receives the incentive signals sent from the upper level. and Subsequently, the core optimization objective is to minimize its overall operating costs (including electricity purchase costs, fuel costs for internal equipment power generation, lifespan depreciation costs, and comfort losses, etc.). The objective function formula is as follows: (19) In the formula, For the first The total operating cost of each lower-level power distribution system; This refers to the comprehensive operating and depreciation costs of its internal heterogeneous distributed resources.

[0062] During this optimization process, the following physical constraints on device operation must be strictly met. First, node energy flow must satisfy power balance constraints: (20) In the formula, and These refer to the actual active power output of the photovoltaic and wind power systems within the system. and These are the total discharge and total charging power of the system's internal energy storage devices (such as batteries or pumped hydro storage); This represents the local raw base load requirement for this node; The amount of controllable loads from industrial or residential sectors that participate in response reduction (or shifting).

[0063] Secondly, for batteries or pumped-storage energy storage devices that participate in internal regulation, charging and discharging mutual exclusion constraints must be met: (twenty one) In the formula and This refers to the charging and discharging power of the energy storage. To ensure the linear solvability of the game model in the solver, binary state variables are introduced in the actual calculation. With the maximum constant (i.e., the Big-M method) transforms this nonlinear constraint into a system of linear inequalities: and Similarly, the aforementioned inter-grid power anti-cycle constraint is also linearized by introducing a power purchase and sale status flag bit.

[0064] For controllable loads in industrial or residential areas, considering their "time shift" property, the load shift conservation constraint must be satisfied: (twenty two) In the formula, This represents the amount of load that is reduced (negative) or replenished (positive) during a specific time period; this global equality constraint ensures the scheduling cycle. The total electricity demand of internal users is not forcibly removed; only time-series reconfiguration is performed. Meanwhile, to ensure a good electricity experience for users' production and daily life, the maximum time span of a single load shift is limited by the maximum delay tolerance time. This parameter Corresponding to the "recovery delay" feature of industrial controllable load extracted in step S101 above, that is, when the algorithm plans and schedules, the absolute value of the time difference between the load reduction action and its corresponding replenishment action must meet the following condition. .

[0065] It should be noted that the actual electrical response boundary calculated based on the second-order attenuation mapping in the preliminary step S2 is... (e.g., limited by battery life degradation) Limited by the mechanical fatigue of the fan (etc.), which are directly and strictly substituted here as inequality constraints (e.g. This mapping mechanism limits the lower-level entities' extreme pursuit of economic interests to the physical safety boundaries defined by equipment lifespan depletion, SOC limits, and mechanical fatigue, thus preventing destructive over-scheduling driven by profit.

[0066] To visually demonstrate the technical effectiveness of the aforementioned secondary decay mapping model, taking battery energy storage systems and wind power generation systems as examples, the three-dimensional fitting relationship between their energy response and dynamic incentive price and response duration is as follows: Figure 2 As shown. From Figure 2 The topological features of the three-dimensional curved surface space can be clearly observed: In the range of low incentive prices (0~200 yuan / MWh), the response capacity of the energy storage system (blue surface) exhibits a rapid, linear sensitivity ramp-up with a high slope; however, when the incentive price exceeds the critical threshold and enters the extreme high-price zone (>500 yuan / MWh) and the continuous call time is prolonged, the response capacity of the battery energy storage system is limited by the response degradation coefficient. The nonlinear penalty effect of the response surface results in a significant “saturation voltage drop” phenomenon; in contrast, the surface of the wind power system (yellow surface) rapidly flattens out and exhibits reverse fine-tuning after reaching the safe limit capacity boundary. Figure 2 The data simulation surface verified the physical reliability of the quadratic response model of the present invention in avoiding overload damage and severe lifespan reduction of the underlying equipment induced by extreme high prices.

[0067] Through the implementation of step S5, the two-layer master-slave game model established by this method effectively coordinates the conflict between the scheduling benefits of the upper-layer global system and the economic interests of the lower-layer local participants, while fully ensuring the physical hard constraints of the operation of the electrical equipment at the bottom layer of the distribution network. By utilizing multi-round dynamic strategy iterations at both the upper and lower layers, the optimal response of the network's flexibility resources under the safety boundary is finally achieved, reaching the optimal Nash equilibrium among all stakeholders.

[0068] Step S6: Solve the game model using an optimization algorithm that introduces dynamic adjustment of inertial weights, and output the optimal dynamic incentive strategy and scheduling execution instructions.

[0069] In this embodiment, the optimization algorithm is the particle swarm optimization algorithm, and the algorithm flow is as follows: Figure 3 As shown. The optimization algorithm that uses inertia weights for dynamic adjustment to solve the game model and output the optimal dynamic incentive strategy and scheduling execution instructions specifically includes the following sub-steps: S601, the time-division energy interaction excitation signal is encoded into the position vector of a particle in the particle swarm and given an initial velocity; Specifically, in the two-level master-slave game model, the dynamic incentive strategy formulated by the upper-level scheduling unit (i.e., the sequence of electricity purchase and sale prices issued to each lower-level distribution system) is a typical high-dimensional continuous decision variable. To achieve efficient optimization, the incentive signals for all time periods within a scheduling cycle are mapped and encoded as spatial coordinate features of the particle swarm optimization algorithm in a high-dimensional search space. Assume the spatial dimension of the electricity price strategy is... ,in (Right now The lower-level power distribution system (the total electricity purchase and sales prices within each time period), then the [number]th [period] [is] [the total number of electricity purchase and sales prices within each time period]. The position vector of each particle This corresponds to a specific time-sharing interactive electricity pricing trial scheme; simultaneously, the flight velocity vector of each particle in the population is randomly initialized. This drives the particles to explore a better price strategy space in subsequent iterations.

[0070] S602, the current particle's position vector is passed as a known parameter to the lower-level follower collaborative optimization model to obtain the optimal interaction power of the bottom-level feedback; Specifically, once a particle's position (i.e., a fixed set of interactive electricity pricing strategies) in the upper-level population is fixedly assigned, the lower-level distribution system uses this price as a known exogenous fixed parameter. In the underlying computing environment, nested calls are made to the YALMIP modeling toolbox under the MATLAB platform, along with mature mathematical programming solver modules such as CPLEX or Gurobi, to solve the lower-level cost minimization model while strictly adhering to physical constraints such as node power balance, energy storage charging and discharging, and load shifting. After the solution is completed, the lower-level distribution system feeds back to the upper level the physical output response plan of its micro-devices and the corresponding optimal interactive electricity results.

[0071] S603, Substitute the optimal interactive power into the objective function of the upper-level leader dynamic incentive model to calculate the current fitness of the particle; After receiving the optimal electricity purchase and sale interaction amounts from all lower-level nodes, the upper-level scheduling unit substitutes these amounts, along with the electricity price strategy corresponding to the current particle, into the upper-level total revenue objective function constructed in step S5 above for calculation. The calculated global net profit value serves as the current fitness, which quantifies the merits of the current particle's position. A higher fitness value indicates that the electricity price strategy represented by the particle is more effective in coordinating the conflict of interests between the upper and lower levels, maximizing the upper-level revenue.

[0072] S604, update the particle's velocity and position vectors according to the current fitness until the maximum number of iterations is reached, and output the optimal dynamic incentive strategy and scheduling execution instruction.

[0073] After evaluating the current fitness of all particles in the population, the algorithm records and updates the optimal coordinates in the search space. Further, updating the particle's velocity and position vectors based on the current fitness includes: S605 utilizes dynamically adjusted inertial weights. Combining the individual optimal position and the global optimal position, the particle's velocity and position vectors are updated using the following formula: (twenty three) (twenty four) in, and The first During the nth iteration The search speed and current position of each particle; and The first During the nth iteration The velocity and position of each particle; For the first The optimal position found by each particle so far; This is the globally optimal position found so far for the entire particle population; and Acceleration factors for individual particles and populations; and These are random numbers distributed between [0,1], used to increase the random perturbation characteristics of the population's flight trajectory.

[0074] To prevent particles from flying out of the effective search space during iteration, the particle velocity and position need to be limited after each update: if Exceeding the preset maximum flight speed If so, then truncate it to the boundary value; In one alternative implementation, for updating the trial electricity price (i.e., particle position), updating the particle's velocity and position vectors based on the current fitness further includes: when the updated position... Exceeding the set lower limit of the main network-side benchmark incentive or upper limit At this time, the conventional rigid boundary truncation strategy is abandoned, and instead an asymmetric elastic reflection over-limit repair mechanism is triggered for position correction, specifically including: Determine the updated position Does it exceed the set lower limit of the main network-side benchmark incentive? or upper limit If so, calculate the deviation from the limit. ,in To determine the value at the corresponding effective boundary, an elastic reflection coefficient following a log-normal distribution is introduced. The excessive trial electricity price can be bounced back into the feasible region using the following formula: (25) in, The corrected particle position. This represents the current iteration number. This represents the maximum number of iterations; the formula takes a plus sign when the lower limit is exceeded, and a minus sign when the upper limit is exceeded.

[0075] When solving the high-dimensional continuous decision space of a multi-level master-slave game model, a rigid truncation method of directly "pulling back and adsorbing to the boundary" for out-of-bounds particles can lead to a large accumulation of high-quality particles at the edge of the search interface in the later stages of iteration, causing a severe "collapse" of population diversity. This makes the algorithm prone to getting trapped in local optima (i.e., unable to derive the optimal purchase and sale price that truly balances the interests of multiple stakeholders). Therefore, during module computation, the deviation degree of the trial price from the limit is extracted first. Subsequently, the elastic reflection coefficient based on the log-normal distribution was used. (By utilizing its right-skewed distribution characteristics, it ensures that the reflection has random perturbations while avoiding extremely large bounce step sizes) and the iterative depth decay term. Constructing dynamic damping. In the initial and exploratory phases of the algorithm search ( With a relatively small decay term close to 1, the out-of-bounds particles are given a large elastic rebound depth, asymmetrically ejecting them into a wide feasible region, effectively maintaining the initial flight inertia of the particles and the exploration diversity of the entire population; in the later stages of algorithm convergence ( Approaching (The decay term is close to 0), and the rebound step size weakens exponentially, prompting particles to perform high-precision small neighborhood optimization close to the system boundary, preventing high-frequency ineffective oscillations from occurring near the boundary.

[0076] Conventional optimization algorithms often employ a rigid truncation strategy, forcibly pulling outwards from the boundary value, when dealing with out-of-bounds particles. This leads to a large accumulation of particles at the search boundary in the later stages of iteration. This diminishes the algorithm's exploration capability, making it difficult to escape local minima. This method addresses the loss of population diversity caused by the limited search space in high-dimensional non-convex game models. It utilizes an asymmetric elastic reflection mechanism to bounce out-of-bounds particles back into the feasible region with dynamic damping. By introducing an elastic coefficient that decays with iteration depth, it maintains particle exploration diversity in the early stages of optimization and achieves refined convergence in the later stages, effectively solving the problem of global optimal solution search accuracy in multi-agent conflict environments.

[0077] As a preferred implementation, the inertia weight in the formula By introducing a progression with iteration number The evolution evolved into a linearly decreasing dynamic adjustment mechanism, employing .

[0078] In a preferred parameter configuration, to ensure that the algorithm balances convergence speed and global optimization capability, the maximum inertia weight... The range of values ​​is Minimum inertia weight The range of values ​​is Acceleration factor and Set all to At the same time, the particle population size is set to... Maximum number of iterations for .

[0079] In the early stages of the algorithm search Maintaining a larger value endows particles with strong kinetic energy to sustain initial flight inertia, which significantly expands the algorithm's exploration range in the global space, helping the population escape local extremum traps; in the later stages of the algorithm's search, Gradually and dynamically decaying to a smaller value reduces flight inertia, prompting the particle to approach the global optimum. It performs refined, localized, and precise convergence mining within a small neighborhood.

[0080] The above iterative process continues until the specified number of iterations is reached. Reaching the preset maximum number of iterations The algorithm determines that it has converged and terminates its operation. At this point, the output of the globally optimal solution is the optimal Nash equilibrium solution of the game model. Based on this, the upper-level scheduling unit issues the optimal dynamic incentive strategy downwards, which in this case is the electricity price strategy, and triggers the underlying scheduling execution instructions.

[0081] Through the implementation of step S6, this method elucidates and deploys a nested optimization mechanism based on improved PSO, effectively overcoming the underlying technical pain point that large-scale non-convex two-layer master-slave game models are difficult to solve analytically due to the large number of nonlinear physical constraints contained in the underlying power grid. The dynamic weight adjustment strategy balances global space exploration and local high-speed convergence, effectively ensuring the computational efficiency and convergence reliability of high-dimensional complex optimization scheduling models in practical engineering applications.

[0082] In simulation tests based on the IEEE 33-bus distribution system architecture, the convergence curves of the improved particle swarm optimization algorithm used to solve the two-layer master-slave game model are compared as follows: Figure 4 As shown. Figure 4 This demonstrates how the cumulative net profit objective function value of the upper-level scheduling unit changes with the number of iterations within a multi-dimensional price optimization space. The evolutionary trajectory. By comparing with the traditional particle swarm optimization algorithm (red dashed line in the figure), it can be seen that the improved particle swarm optimization algorithm (green solid line in the figure) has a higher fitness improvement rate in the early stages of iteration (0-50 iterations); in the middle and later stages of iteration (after 150 iterations), due to the introduction of a method based on... The algorithm employs asymmetric heuristic elastic reflection logic, achieving non-rigid correction of the position vector at the search space boundary. Figure 4The convergence trajectory shows that the improved particle swarm optimization algorithm, through the synergistic effect of dynamic weights and elastic reflection mechanism, enhances the exploration capability of the population in the later stages of the search space, reduces the probability of premature convergence in the non-convex optimization process, and improves the search accuracy of the global optimal solution.

[0083] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0084] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0085] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0086] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0087] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0088] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A distributed resource optimization scheduling method for flexibility improvement, characterized in that, include: Obtain the multidimensional feature vectors and adjustment strengths of heterogeneous distributed resources, and construct a dynamic activation function; A quadratic decay mapping model is constructed based on the dynamic excitation function to quantify the electrical energy response boundary. The latent features of the multidimensional feature vector are extracted by a deep autoencoder, and the heterogeneous distributed resources are divided into multiple virtual machine groups by combining the distribution network topology constraints. A two-layer master-slave game model is established between the upper-layer scheduling unit and the lower-layer power distribution system based on the virtual machine group, wherein the power response boundary serves as the physical constraint condition for each virtual machine group in the lower-layer power distribution system to respond to the upper-layer scheduling command. The game model is solved using an optimization algorithm that incorporates dynamic adjustment of inertia weights, and the optimal dynamic incentive strategy and scheduling execution instructions are output.

2. The method of claim 1, wherein, The quadratic decay mapping model includes constructing a response model for distributed resources, including various nonlinear decay characteristics and boundary capacity constraints, by using the output of the dynamic excitation function as the independent variable. The distributed resources include battery energy storage systems, wind power generation systems, and controllable industrial loads. The dynamic excitation function is expressed as: in, The coefficient of ascent speed. The maximum incentive price cap, For inflection point parameters, To adjust for intensity variables, where .

3. The method according to claim 1, characterized in that, The virtual machine group is divided as follows: Based on the aforementioned latent features, a generalized clustering distance metric for integrating distribution network topology constraints is constructed, and clustering is performed accordingly. The generalized clustering distance measure includes the potential feature distance between resources, the electrical impedance penalty term constructed based on the distribution network impedance, and the exponential penalty term constructed based on the node voltage sensitivity matrix, the maximum adjustable output boundary and the node voltage critical value. Clustering algorithms are used to aggregate clusters based on the generalized clustering distance metric. Aggregation stops when the generalized clustering distance of the merged clusters is greater than the set truncation threshold, and multiple virtual machine groups are created.

4. The method according to claim 1, characterized in that, The two-layer master-slave game model executes a dynamic incentive strategy by constructing a multi-timescale hierarchical scheduling architecture that includes day-ahead, intraday, and real-time timescales. The dynamic incentive strategy includes: During the current scheduling phase, a two-stage stochastic planning based on scenario analysis is used to formulate the basic power plan for each virtual machine group; During the intraday rolling scheduling phase, model predictive control is adopted to track energy interaction benchmark results while satisfying the linearized power flow constraints of the distribution network. During the real-time control phase, when a sudden fault or frequency voltage exceedance is detected in the distribution network, a high-sensitivity node is located using a sensitivity matrix, and the corresponding heterogeneous distributed resources of the high-sensitivity node are controlled to inject reactive power to provide voltage support and perform a voltage exceedance response.

5. The method according to claim 1, characterized in that, The establishment of a two-layer master-slave game model between the upper-layer scheduling unit and the lower-layer power distribution system includes: A dynamic incentive model for upper-level leaders is constructed, with the maximization of scheduling benefits as the objective function. Time-sharing energy interaction incentive signals are formulated under the constraints of inter-network power anti-cyclic constraints and the main grid-side benchmark incentive boundary. A lower-level follower collaborative optimization model is constructed, with the goal of minimizing operating costs. Under the constraints of node power balance, load shift conservation, and mutual exclusion of energy storage system charging and discharging, the output of internal equipment and the amount of electricity exchanged are optimized.

6. The method according to claim 5, characterized in that, The optimization algorithm is a particle swarm optimization algorithm; the optimization algorithm that uses dynamically adjusted inertia weights to solve the game model and outputs the optimal dynamic incentive strategy and scheduling execution instructions includes: The time-division energy interaction excitation signal is encoded into the position vector of a particle in the particle swarm and assigned an initial flight velocity; The current particle's position vector is passed as a known parameter to the lower-level follower collaborative optimization model to obtain the optimal interaction power of the bottom-level feedback; Substitute the optimal interaction power into the objective function of the upper-level leader dynamic incentive model to calculate the current fitness of the particle. Update the particle's velocity and position vectors based on the current fitness until the maximum number of iterations is reached, and then output the optimal dynamic incentive strategy and scheduling execution instructions.

7. The method according to claim 6, characterized in that, The step of updating the particle's velocity and position vectors based on the current fitness includes: Using dynamically adjusted inertia weights, and combining the individual optimal position and the global optimal position, the particle's velocity and position vectors are updated using the following formula: in, and The first During the nth iteration The velocity and position of each particle and The first During the nth iteration The velocity and position of each particle For the optimal position of an individual, To be the globally optimal position and Acceleration factors for individual particles and populations, and It is a random number. This is the inertial weight.

8. The method according to claim 4, characterized in that, The control of injecting reactive power into the heterogeneous distributed resources corresponding to the high-sensitivity node to provide voltage support includes: The heterogeneous distributed resources corresponding to the high-sensitivity nodes are used as grid-type inverters, and reactive power is adaptively injected using a droop control law based on reactive capacity margin adaptive adjustment. The control law is expressed as follows: in, For nodes The amount of reactive power injected in an emergency to support the grid-type inverter. Based on the droop control coefficient For nodes The current actual reactive power generated by the grid-connected inverter. To constrain the physical boundary of the maximum available reactive power due to the equipment's rated parameters, As a margin penalty adjustment factor, For nodes Reference rated voltage, For real-time voltage measurement.

9. The method according to claim 7, characterized in that, The step of updating the particle's velocity and position vectors based on the current fitness further includes: When the updated position exceeds the set lower or upper limit of the main network side reference excitation, the asymmetric elastic reflection over-limit repair mechanism is triggered to correct the position.

10. The method according to claim 9, characterized in that, The position correction includes: Calculate the deviation from the limit , in, The updated position; The value at the corresponding valid boundary where the breakthrough occurs; Introducing an elastic reflection coefficient that follows a log-normal distribution, the excessive trial electricity price is bounced back into the feasible region using the following formula: in, The corrected particle position. This represents the current iteration number. The maximum number of iterations, is the elastic reflection coefficient.