Distribution network auxiliary decision-making method and system considering source load fluctuation relevance, and medium

By building a deep perception mechanism for aggregated information in the station area and strengthening the learning decision-making structure, quantifying the correlation between source and load fluctuations, the decision-making problems of traditional distribution networks under high proportion of distributed energy and diversified loads are solved, efficient, precise decision-making and risk management of distribution networks are achieved, and the safe and economic operation of the new power system is improved.

CN120579842APending Publication Date: 2025-09-02SUQIAN POWER SUPPLY COMPANY OF JIANGSU PROVINCE POWER +2
View PDF 0 Cites 26 Cited by

Patent Information

Application Number
CN202510653724.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When traditional distribution networks face strong random coupling of high proportion distributed energy and multi-load, they cannot achieve global optimal decisions. The existing methods are difficult to cope with the spatial and temporal correlation characteristics of source load minute-level fluctuations, resulting in multi-dimensional and cross-scale transmission of operation risks, and the decision information is not refined enough, and traditional optimization algorithms are difficult to cope with real-time fluctuations. The existing reinforcement learning methods have poor multi-objective collaboration and poor strategy interpretability.

Method used

Build a deep perception mechanism for aggregated information in the station area, design a reinforcement learning decision-making architecture, quantify source-charge fluctuations correlation, establish a composite feature vector, build a deep reinforcement learning model, use transfer learning to decouple scene features, realize the independent evolution and rapid adaptation of strategies, and build a power grid digital twin simulation platform for strategy verification.

Benefits of technology

It improves the decision-making efficiency and accuracy of the distribution network, enhances the adaptability to new scenarios, realizes the accurate perception and response of distribution network operation risks, and improves the safe and economic operation level of the new power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention relates to the technical field of power systems and automation thereof, in particular to a distribution network auxiliary decision-making method and system considering source load fluctuation relevance and a medium. The method comprises the following steps: firstly, collecting related information of a distribution network area, quantifying a synchronization and hysteresis association rule of multi-source heterogeneous data fluctuation, and constructing a composite feature vector and a standardized risk perception data set; defining a state space and an action space of a reinforcement learning algorithm based on the composite feature vector, and realizing auxiliary decision-making optimization of the distribution network; constructing a scene feature library, calculating the fluctuation relevance similarity between a new scene and a historical scene, and multiplexing a deep reinforcement learning model architecture and carrying out transfer learning; building a power grid digital twinborn simulation platform, designing evaluation indexes, generating candidate schemes, deducing the candidate schemes, selecting recommendation strategies and storing the recommendation strategies in a strategy knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems and automation technologies thereof, and in particular to a distribution network auxiliary decision-making method considering source-load fluctuation correlation. Background Art

[0002] Against the backdrop of new power system construction, the high proportion of distributed energy resources and the strong stochastic coupling of diverse loads are driving a profound shift in distribution network operation from "source follows load" to "source and load interact." This shift poses significant challenges to traditional decision-making systems based on simple risk perception methods. The spatiotemporal correlations of minute-by-minute fluctuations in sources and loads lead to multi-dimensional, cross-scale transmission of operational risks in distribution networks. Previous strategies using single-device control or local optimization cannot achieve global optimality. Building a decision-support system adaptable to dynamic scenarios is crucial for resolving the contradiction between efficient renewable energy integration and power supply reliability. It is also key to promoting the development of proactive and intelligent distribution networks and holds significant engineering value for the safe and economical operation of new power systems. Current research reveals several key limitations. In terms of data collection, there is insufficient in-depth aggregation of substation-level device status and distributed resource output, resulting in isolated multi-source information and insufficiently refined decision-making information. Decision optimization relies on offline simulation and manual experience, making traditional optimization algorithms incapable of coordinating multiple objectives and providing poor policy interpretability. In the face of new scenarios, there is a lack of effective dynamic migration mechanisms, and decision-making models often fail due to shifts in data distribution.

[0003] To overcome these bottlenecks, a new, systematic solution for distribution network decision-making assistance is urgently needed. This invention pioneers a deep perception mechanism for aggregated information across substations, enabling cross-scale correlation analysis between device status and grid behavior. It then designs a reinforcement learning decision-making architecture to transform operational risk prevention and control objectives into an autonomously evolving strategy space. Finally, it utilizes transfer learning to decouple scenario characteristics and enable widespread application of strategy knowledge. This innovative solution is expected to bring new breakthroughs to distribution network decision-making assistance and significantly advance the development of new power systems. Summary of the Invention

[0004] This invention aims to overcome the limitations of traditional distribution networks based on simple risk perception and proposes a distribution network decision-making support method that considers the correlation between source and load fluctuations. This method quantifies the correlation between fluctuations at a probabilistic level, resolving the feature distortion problem caused by traditional independent modeling and laying a high-fidelity data foundation for distribution network decision-making support. Furthermore, a composite reward function is designed to simultaneously constrain immediate prediction accuracy and long-term early warning capabilities, overcoming the short-sightedness of supervised learning models in time-series decision-making.

[0005] A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations includes the following steps:

[0006] S1: Collect distribution network substation information, quantify the correlation between the synchronization and hysteresis of multi-source heterogeneous data in PV output and load power fluctuations, calculate the composite feature vector based on this correlation, and establish a standardized risk perception dataset containing aggregated information of distribution network substations;

[0007] S2: Define the state space and action space of the reinforcement learning algorithm based on the composite feature vector, design the reward function, and build a deep reinforcement learning model architecture to assist in decision-making and optimization of the distribution network.

[0008] S3: Build a scene feature library. For a new target scene, calculate its fluctuation correlation similarity with historical scenes in the feature library, select the reference scene set with the strongest correlation, use the learning algorithm to train the deep reinforcement learning model architecture, freeze the feature extraction layer parameters, and retain its ability to analyze source-load correlation features. Use a composite reward function to constrain the transfer training direction to achieve rapid scene adaptation.

[0009] S4: Build a digital twin simulation platform for the power grid and design evaluation indicators in different dimensions. Input the real-time distribution network substation dataset into the migration-trained deep reinforcement learning model architecture to generate several groups of differentiated candidate solutions. Use the digital twin simulation platform to deduce all solutions and compare some of them to obtain recommended strategies. After successful verification, the recommended strategies will be stored in the strategy knowledge base. When there are similar scenarios, the strategies will be executed first.

[0010] Preferably, the step S1 includes the following sub-steps:

[0011] S101. Multi-source heterogeneous data collection and cleaning: Collect distribution network area related information, including distribution network area aggregated information, distributed power output time series data, load power curves, meteorological information and network topology parameters, historical limit-crossing event records, and environmental variables for the corresponding time period. Use the quartile method to eliminate outliers in the distribution network area related information, and then use the sliding window mean interpolation and adjacent time period similarity matching algorithm to repair the missing data.

[0012] S102. Data alignment and feature decoupling in spatiotemporal dimensions: Interpolation alignment is performed on the collected multi-source heterogeneous data based on a unified 5-minute time resolution. Geographic information matching technology is used to correlate the load data of distributed power generation nodes and power supply areas. At the same time, a sliding time window is used to segment the time series data, and statistical features of source-load fluctuations within each time window are extracted, including but not limited to extreme values, variance, and volatility. Subsequently, the fluctuation components of weather-sensitive loads and rigid loads are separated based on the correlation between weather changes and load changes.

[0013] S103. Quantitative modeling of fluctuation correlation characteristics: The grey correlation analysis method is used to calculate the correlation index between the source-side output fluctuation and the load-side power change at different time scales. A fluctuation pattern matching library is constructed based on the Hilbert-Huang transform to identify the coupling characteristics of source-load fluctuations in the time-frequency domain. A joint probability distribution model of source-load fluctuations is established through the non-parametric kernel density estimation method to quantify the synchronization and lag correlation between the two fluctuations.

[0014] S104. Composite feature vector construction and label definition: The raw data features including PV output and load power are integrated with the fluctuation synchronization coefficient and the fluctuation hysteresis coefficient in multiple dimensions to obtain a composite feature vector including timestamp, spatial location, fluctuation pattern, and correlation strength. Risk level labels are then defined based on the deviation between the actual power of the line and the capacity threshold. Dynamic threshold correction technology is used to eliminate the influence of seasonal characteristics on label distribution.

[0015] S105. Scenario-based dataset division and enhancement: The dataset is classified into scenarios according to meteorological type, load characteristics, and network topology. Then, for the scarce risk scenario samples and small sample scenarios after classification, the SMOTE oversampling algorithm and the virtual sample generation strategy based on the adversarial generative network are used to perform data enhancement. Finally, the validity of the associated features is verified to obtain a standardized risk perception dataset with strong representation capabilities.

[0016] Preferably, in step S103, the synchronicity of the fluctuations is expressed as:

[0017]

[0018] Where I(X,Y) is the synchronization coefficient of the discretized observations of the source-side power fluctuation sequence X and the load-side power fluctuation sequence Y; p(x,y) is the joint probability distribution of element x in X and element y in Y; p(x) and p(y) are the marginal probability distributions of X and Y.

[0019] The hysteresis of the fluctuation is expressed as:

[0020]

[0021] Among them, τ is the time lag parameter; R XY is the hysteresis coefficient of the source side power fluctuation sequence X and the load side power fluctuation sequence Y; t ,Y t+τ are the fluctuation values ​​of the source side at time t and the fluctuation values ​​of the load side at time t+τ respectively; μ X ,μ Y are the means of the source-load fluctuation series respectively; t is the time value; N is the total number of data points in the source-load side power fluctuation time series.

[0022] Preferably, in the step S205, when verifying the validity of the associated features: a feature contribution analysis method based on Shapley value is used to quantify the risk perception contribution weight of each associated feature to the over-limit risk, and by comparing the difference in model risk perception performance of the actual power grid operation data with and without associated features, redundant features are eliminated and the feature combination is optimized.

[0023] Preferably, the step S2 includes the following sub-steps:

[0024] S201, dynamic state space construction and feature encoding: Based on the composite feature vector generated in step S1, the state space of the reinforcement learning algorithm is defined. The state vectors added include but are not limited to: distribution network area aggregation information, real-time source-load fluctuation correlation indicators, historical line limit-crossing frequency, current environmental variables, and topological structure characteristics;

[0025] S202, Action Space Design and Risk Level Mapping: The action space of the reinforcement learning algorithm is defined by pre-delineating the probability intervals for network operation risk perception. The intervals cover low, medium, and high risk levels, each corresponding to a warning level under a specific probability threshold. An adaptive action classification mechanism is introduced to dynamically adjust the interval boundaries based on the line load factor.

[0026] S203. Deep reinforcement learning network architecture implementation: Build a hybrid neural network based on the LSTM-AC framework, deploy a dynamic experience pool to store interaction data, and use a priority sampling mechanism to give a higher selection probability to training samples in high-risk periods;

[0027] S204. Deploy a dynamic experience pool to store interaction data, and use a priority sampling mechanism to give a higher selection probability to training samples in high-risk periods.

[0028] Preferably, during the reinforcement learning, a reward function integrating multi-objective optimization is designed, which can be specifically expressed as:

[0029]

[0030] Among them, R t The total reward value at time t in the direction of guiding strategy optimization; is an indicator function, when the risk perception state a t and the actual limit-crossing state y t If they are consistent, it is 1, otherwise it is 0; ω1, ω2, ω3 are weight coefficients; P t pred is the predicted line power at time t; P t actual is the actual line power at time t; P th is the line safety power threshold; K is the look-ahead time step

[0031] Preferably, in step S204, LSTM-AC is a long short-term memory-executor evaluator, and the executor network in the LSTM-AC framework adopts a bidirectional LSTM structure to capture the temporal correlation characteristics of source-load fluctuations and output action probability distribution.

[0032] Preferably, the step S3 includes the following sub-steps:

[0033] S301. Scenario feature library construction: Based on the standardized data set established in step S2, source-load fluctuation correlation characteristics and environmental parameters for various scenarios are extracted. Scenarios include, but are not limited to, high photovoltaic power generation and sudden load changes. Source-load fluctuation correlation characteristics include, but are not limited to, synchronization coefficients and lagged correlation coefficients. Based on these data, a multi-dimensional scenario feature library is constructed, including fluctuation patterns, network topology, and meteorological conditions, as a basic carrier of transferable knowledge.

[0034] S302. Target scenario feature similarity assessment: For the new target scenario, calculate its fluctuation correlation similarity with the historical scenarios in the feature library, quantify the consistency of the source-load fluctuation pattern, and simultaneously select the reference scenario set with the strongest correlation as the data basis for migration;

[0035] S303, scene migration: Reuse the deep reinforcement learning model architecture trained in step S2, freeze the feature extraction layer parameters, retain its ability to analyze source-load correlation features, fine-tune only the fully connected layer of the policy network for a small amount of labeled data of the target scene, and use the composite reward function designed in step S2 to constrain the migration training direction to achieve rapid scene adaptation.

[0036] Preferably, the similarity evaluation model between the target new scene and the source scene is specifically:

[0037]

[0038] Among them, D KL is the distribution difference measure between the source scene and the target scene; P(x) is the joint probability distribution of the source scene; Q(x) is the joint probability distribution of the target scene, that is, the estimated value of the new scene; The feature space includes but is not limited to source-load correlation indicators and fluctuation patterns.

[0039] Preferably, the step S4 includes the following sub-steps:

[0040] S401. Construction of a simulation environment for auxiliary decision-making strategies: Based on the source-load fluctuation correlation dataset established in step S1 and the auxiliary decision-making model trained in step S2, a power grid digital twin simulation platform is constructed to support dynamic deduction of auxiliary decision-making strategies.

[0041] S402. Design of a multi-dimensional evaluation index system: Design evaluation indicators for economy, safety, and renewable energy absorption rate. Economy includes, but is not limited to, network losses and dispatch instruction execution costs; safety includes, but is not limited to, over-limit risk attenuation rate and voltage compliance rate; and renewable energy utilization includes, but is not limited to, curtailed solar power rate, curtailed wind power rate, and load matching.

[0042] S403, Intelligent Generation of Candidate Scheduling Strategies: In combination with the transfer learning model from step S3, a candidate strategy set is automatically generated based on the current over-limit risk level, including but not limited to: energy storage charging and discharging power optimization, interruptible load regulation, and distributed power generation output control. 10-15 sets of differentiated candidate solutions are then generated.

[0043] S404. Dynamic simulation verification of strategy effects: Utilize the grid digital twin simulation platform established in step S401 to calculate all differentiated candidate solutions generated in step S403, and collect the time series data of key indicators in the calculation, as well as the over-limit risk value, smoothness of the new energy consumption curve, and node voltage fluctuation rate within 1 hour after the implementation of the solution, and then compare and score them with the undispatched benchmark scenario.

[0044] S405, Comprehensive Comparison and Optimization of Multiple Strategies: Evaluate the long-term benefits of each strategy based on the subsequent risk evolution trend of the decision-making support to set the strategy validity period attenuation factor. Select the three most consistent strategies with significant short-term benefits and controllable long-term negative impacts from the 0-15 groups of differentiated candidate solutions generated in S403 as recommended strategies and generate a quantitative benefit comparison table.

[0045] S406. Dynamic evaluation report generation and feedback: Based on the recommended strategy and the quantitative benefit comparison table in S405, an evaluation report including a strategy execution plan, an expected effect map, and risk warning information is generated and presented to the operator, and the verification results of the recommended strategy are awaited;

[0046] S407, iterative update of the strategy knowledge base: the verified effective recommended strategies are classified and stored in the strategy knowledge base according to the scenario characteristics. When a similar scenario triggers an out-of-limit warning, the strategy is executed first or the parameters are fine-tuned based on the real-time effect.

[0047] A distribution network auxiliary decision-making system considering the correlation between source and load fluctuations is characterized by comprising a perception layer, a decision layer, a migration layer, and a verification layer;

[0048] The perception layer integrates distribution network area aggregation information and area-level edge computing nodes to collect real-time output fluctuations of distributed power sources, distribution transformer load rates, and user-side response potential;

[0049] The decision-making layer is deployed with a cloud-edge collaborative reinforcement learning decision engine to generate an optimization strategy that considers the spatiotemporal correlation between source and load;

[0050] The migration layer is used to establish a scene feature knowledge base to achieve safe migration of decision strategies in different penetration scenarios;

[0051] The verification layer is used to build a digital twin verification platform to implement strategy preview and effect tracing.

[0052] Preferably, the closed-loop operation mechanism includes a "dynamic perception-decision closed loop" and a "cross-scenario migration closed loop." The dynamic perception-decision closed loop begins with data collection and completes closed-loop iterations through online strategy generation, instruction issuance and execution, and effect feedback. The cross-scenario migration closed loop begins with new scenario feature extraction and completes closed-loop iterations through strategy parameter migration, virtual simulation verification, and incremental knowledge base updates.

[0053] A computer-readable storage medium, characterized in that a computer program is stored thereon, and when the computer program is executed, a distribution network voltage control method under uncertain factors according to any one of claims 1 to 10 is implemented.

[0054] The beneficial effects of the present invention are:

[0055] 1. This invention collects and cleans multi-source heterogeneous data, uses the quartile method to eliminate outliers, and employs a sliding window mean interpolation and adjacent time period similarity matching algorithm to repair missing data, ensuring data accuracy and integrity. Geographic information matching technology is used to correlate distributed power supply nodes with power supply area load data, performing interpolation alignment operations. A sliding time window is used to extract the statistical characteristics of source-load fluctuations, separating meteorologically sensitive loads from rigid load fluctuation components. The resulting standardized risk perception dataset has strong characterization capabilities, providing a high-quality data foundation for subsequent decision-making and significantly improving the accuracy of understanding source-load fluctuation patterns.

[0056] 2. This paper constructs a dynamic state space and design action space based on composite feature vectors, and uses a reward function integrated with multi-objective optimization to guide strategy optimization. By deploying a hybrid neural network based on the LSTM-AC framework, combined with dynamic experience pool storage and a priority sampling mechanism, the reinforcement learning algorithm can effectively capture the temporal correlation characteristics of source-load fluctuations, enabling rapid and accurate auxiliary decision-making and optimization in complex distribution network operation scenarios, significantly improving decision-making efficiency and accuracy, and enhancing the ability to perceive and respond to distribution network operation risks.

[0057] 3. This invention constructs a multidimensional scenario feature library. For new target scenarios, it screens a set of reference scenarios by calculating the similarity of fluctuation correlations. It reuses the trained model architecture and freezes the parameters of the feature extraction layer, fine-tuning only the fully connected layer of the policy network. This, combined with a composite reward function, enables rapid scenario adaptation. This ensures that the decision model remains effective across different scenarios, greatly enhancing the model's adaptability to new scenarios, broadening the application scope of distribution network decision-making assistance methods, and reducing the cost and time of model training in new scenarios.

[0058] 4. The present invention builds a digital twin simulation platform for power grids, designs multi-dimensional evaluation indicators such as economy, safety, and new energy absorption rate, combines the transfer learning model to generate differentiated candidate solutions and conduct dynamic simulation verification. Through comprehensive comparison and optimization of multiple strategies, solutions with significant short-term effects and controllable long-term negative impacts are selected as recommended strategies, an evaluation report is generated, and the strategy knowledge base is iteratively updated. This series of operations realizes the comprehensive evaluation and optimization of decision-making strategies, ensures the safety, economy and efficient absorption of new energy in the operation of the distribution network, improves the overall operation level of the distribution network, and helps the safe and economical operation of the new power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The figure is a schematic diagram of the steps of a distribution network auxiliary decision-making method considering the correlation between source and load fluctuations of the present invention.

[0060] Figure 2 This is a risk perception effect diagram under the source-load fluctuation scenario of the present invention.

[0061] Figure 3 This is the effect diagram of the distribution network auxiliary decision-making of the present invention. DETAILED DESCRIPTION

[0062] The present invention aims to break through the limitations of traditional distribution networks based on simple risk perception and proposes a distribution network auxiliary decision-making method that considers the correlation between source and load fluctuations, including the following steps:

[0063] S1: Collect distribution network substation information, quantify the correlation between the synchronization and hysteresis of multi-source heterogeneous data in PV output and load power fluctuations, calculate the composite feature vector based on this correlation, and establish a standardized risk perception dataset containing aggregated information of distribution network substations;

[0064] S2: Define the state space and action space of the reinforcement learning algorithm based on the composite feature vector, design the reward function, and build a deep reinforcement learning model architecture to assist in decision-making and optimization of the distribution network.

[0065] S3: Build a scene feature library. For a new target scene, calculate its fluctuation correlation similarity with historical scenes in the feature library, select the reference scene set with the strongest correlation, use the learning algorithm to train the deep reinforcement learning model architecture, freeze the feature extraction layer parameters, and retain its ability to analyze source-load correlation features. Use a composite reward function to constrain the transfer training direction to achieve rapid scene adaptation.

[0066] S4: Build a digital twin simulation platform for the power grid and design evaluation indicators in different dimensions. Input the real-time distribution network substation dataset into the migration-trained deep reinforcement learning model architecture to generate several groups of differentiated candidate solutions. Use the digital twin simulation platform to deduce all solutions and compare some of them to obtain recommended strategies. After successful verification, the recommended strategies will be stored in the strategy knowledge base. When there are similar scenarios, the strategies will be executed first.

[0067] Step S1 includes the following sub-steps:

[0068] S101. Multi-source heterogeneous data collection and cleaning: Collect distribution network area related information, including distribution network area aggregated information, distributed power output time series data, load power curves, meteorological information and network topology parameters, historical limit-crossing event records, and environmental variables for the corresponding time period. Use the quartile method to eliminate outliers in the distribution network area related information, and then use the sliding window mean interpolation and adjacent time period similarity matching algorithm to repair the missing data.

[0069] S102. Data alignment and feature decoupling in spatiotemporal dimensions: Interpolation alignment is performed on the collected multi-source heterogeneous data based on a unified 5-minute time resolution. Geographic information matching technology is used to correlate the load data of distributed power generation nodes and power supply areas. At the same time, a sliding time window is used to segment the time series data, and statistical features of source-load fluctuations within each time window are extracted, including but not limited to extreme values, variance, and volatility. Subsequently, the fluctuation components of weather-sensitive loads and rigid loads are separated based on the correlation between weather changes and load changes.

[0070] S103. Quantitative modeling of fluctuation correlation characteristics: The grey correlation analysis method is used to calculate the correlation index between the source-side output fluctuation and the load-side power change at different time scales. A fluctuation pattern matching library is constructed based on the Hilbert-Huang transform to identify the coupling characteristics of source-load fluctuations in the time-frequency domain. A joint probability distribution model of source-load fluctuations is established through the non-parametric kernel density estimation method to quantify the synchronization and lag correlation between the two fluctuations.

[0071] S104. Composite feature vector construction and label definition: The raw data features including PV output and load power are integrated with the fluctuation synchronization coefficient and the fluctuation hysteresis coefficient in multiple dimensions to obtain a composite feature vector including timestamp, spatial location, fluctuation pattern, and correlation strength. Risk level labels are then defined based on the deviation between the actual power of the line and the capacity threshold. Dynamic threshold correction technology is used to eliminate the influence of seasonal characteristics on label distribution.

[0072] S105. Scenario-based dataset division and enhancement: The dataset is classified into scenarios according to meteorological type, load characteristics, and network topology. Then, for the scarce risk scenario samples and small sample scenarios after classification, the SMOTE oversampling algorithm and the virtual sample generation strategy based on the adversarial generative network are used to perform data enhancement. Finally, the validity of the associated features is verified to obtain a standardized risk perception dataset with strong representation capabilities.

[0073] In step S103, the synchronicity of the fluctuations is expressed as:

[0074]

[0075] Where I(X,Y) is the synchronization coefficient of the discretized observations of the source-side power fluctuation sequence X and the load-side power fluctuation sequence Y; p(x,y) is the joint probability distribution of element x in X and element y in Y; p(x) and p(y) are the marginal probability distributions of X and Y.

[0076] The hysteresis of the fluctuation is expressed as:

[0077]

[0078] Among them, τ is the time lag parameter; R XY is the hysteresis coefficient of the source side power fluctuation sequence X and the load side power fluctuation sequence Y; t ,Y t+τ are the fluctuation values ​​of the source side at time t and the fluctuation values ​​of the load side at time t+τ respectively; μ X ,μ Y are the means of the source-load fluctuation series respectively; t is the time value; N is the total number of data points in the source-load side power fluctuation time series.

[0079] In step S205, when verifying the validity of the associated features: a feature contribution analysis method based on Shapley value is used to quantify the risk perception contribution weight of each associated feature to the over-limit risk. By comparing the difference in model risk perception performance of the actual power grid operation data with and without associated features, redundant features are eliminated and the feature combination is optimized.

[0080] Step S2 includes the following sub-steps:

[0081] S201, dynamic state space construction and feature encoding: Based on the composite feature vector generated in step S1, the state space of the reinforcement learning algorithm is defined. The state vectors added include but are not limited to: distribution network area aggregation information, real-time source-load fluctuation correlation indicators, historical line limit-crossing frequency, current environmental variables, and topological structure characteristics;

[0082] S202, Action Space Design and Risk Level Mapping: The action space of the reinforcement learning algorithm is defined by pre-delineating the probability intervals for network operation risk perception. The intervals cover low, medium, and high risk levels, each corresponding to a warning level under a specific probability threshold. An adaptive action classification mechanism is introduced to dynamically adjust the interval boundaries based on the line load factor.

[0083] S203. Deep reinforcement learning network architecture implementation: Build a hybrid neural network based on the LSTM-AC framework, deploy a dynamic experience pool to store interaction data, and use a priority sampling mechanism to give a higher selection probability to training samples in high-risk periods;

[0084] S204. Deploy a dynamic experience pool to store interaction data, and use a priority sampling mechanism to give a higher selection probability to training samples in high-risk periods.

[0085] When strengthening learning, the reward function that integrates multi-objective optimization is designed, which can be specifically expressed as:

[0086]

[0087] Among them, R t The total reward value at time t in the direction of guiding strategy optimization; is an indicator function, when the risk perception state a t and the actual limit-crossing state y t If they are consistent, it is 1, otherwise it is 0; ω1, ω2, ω3 are weight coefficients; P t pred is the predicted line power at time t; P t actual is the actual line power at time t; P th is the line safety power threshold; K is the look-ahead time step

[0088] In step S204, LSTM-AC stands for Long Short-Term Memory-Actor Evaluator. The actor network in the LSTM-AC framework uses a bidirectional LSTM structure to capture the temporal correlation characteristics of source-load fluctuations and outputs the action probability distribution.

[0089] Step S3 includes the following sub-steps:

[0090] S301. Scenario feature library construction: Based on the standardized data set established in step S2, source-load fluctuation correlation characteristics and environmental parameters for various scenarios are extracted. Scenarios include, but are not limited to, high photovoltaic power generation and sudden load changes. Source-load fluctuation correlation characteristics include, but are not limited to, synchronization coefficients and lagged correlation coefficients. Based on these data, a multi-dimensional scenario feature library is constructed, including fluctuation patterns, network topology, and meteorological conditions, as a basic carrier of transferable knowledge.

[0091] S302. Target scenario feature similarity assessment: For the new target scenario, calculate its fluctuation correlation similarity with the historical scenarios in the feature library, quantify the consistency of the source-load fluctuation pattern, and simultaneously select the reference scenario set with the strongest correlation as the data basis for migration;

[0092] S303, scene migration: Reuse the deep reinforcement learning model architecture trained in step S2, freeze the feature extraction layer parameters, retain its ability to analyze source-load correlation features, fine-tune only the fully connected layer of the policy network for a small amount of labeled data of the target scene, and use the composite reward function designed in step S2 to constrain the migration training direction to achieve rapid scene adaptation.

[0093] The similarity evaluation model between the target new scene and the source scene is as follows:

[0094]

[0095] Among them, D KL is the distribution difference measure between the source scene and the target scene; P(x) is the joint probability distribution of the source scene; Q(x) is the joint probability distribution of the target scene, that is, the estimated value of the new scene; The feature space includes but is not limited to source-load correlation indicators and fluctuation patterns.

[0096] Step S4 includes the following sub-steps:

[0097] S401. Construction of a simulation environment for auxiliary decision-making strategies: Based on the source-load fluctuation correlation dataset established in step S1 and the auxiliary decision-making model trained in step S2, a power grid digital twin simulation platform is constructed to support dynamic deduction of auxiliary decision-making strategies.

[0098] S402. Design of a multi-dimensional evaluation index system: Design evaluation indicators for economy, safety, and renewable energy absorption rate. Economy includes, but is not limited to, network losses and dispatch instruction execution costs; safety includes, but is not limited to, over-limit risk attenuation rate and voltage compliance rate; and renewable energy utilization includes, but is not limited to, curtailed solar power rate, curtailed wind power rate, and load matching.

[0099] S403, Intelligent Generation of Candidate Scheduling Strategies: In combination with the transfer learning model from step S3, a candidate strategy set is automatically generated based on the current over-limit risk level, including but not limited to: energy storage charging and discharging power optimization, interruptible load regulation, and distributed power generation output control. 10-15 sets of differentiated candidate solutions are then generated.

[0100] S404. Dynamic simulation verification of strategy effects: Utilize the grid digital twin simulation platform established in step S401 to calculate all differentiated candidate solutions generated in step S403, and collect the time series data of key indicators in the calculation, as well as the over-limit risk value, smoothness of the new energy consumption curve, and node voltage fluctuation rate within 1 hour after the implementation of the solution, and then compare and score them with the undispatched benchmark scenario.

[0101] S405, Comprehensive Comparison and Optimization of Multiple Strategies: Evaluate the long-term benefits of each strategy based on the subsequent risk evolution trend of the decision-making support to set the strategy validity period attenuation factor. Select the three most consistent strategies with significant short-term benefits and controllable long-term negative impacts from the 0-15 groups of differentiated candidate solutions generated in S403 as recommended strategies and generate a quantitative benefit comparison table.

[0102] S406. Dynamic evaluation report generation and feedback: Based on the recommended strategy and the quantitative benefit comparison table in S405, an evaluation report including a strategy execution plan, an expected effect map, and risk warning information is generated and presented to the operator, and the verification results of the recommended strategy are awaited;

[0103] S407, iterative update of the strategy knowledge base: the verified effective recommended strategies are classified and stored in the strategy knowledge base according to the scenario characteristics. When a similar scenario triggers an out-of-limit warning, the strategy is executed first or the parameters are fine-tuned based on the real-time effect.

[0104] A distribution network auxiliary decision-making system that considers the correlation between source and load fluctuations includes a perception layer, a decision layer, a migration layer, and a verification layer;

[0105] The perception layer integrates distribution network area aggregation information and area-level edge computing nodes to collect real-time data on distributed power output fluctuations, distribution transformer load rates, and user-side response potential.

[0106] The decision-making layer deploys a cloud-edge collaborative reinforcement learning decision engine to generate optimization strategies that consider the spatiotemporal correlation between sources and loads.

[0107] The migration layer is used to build a scenario feature knowledge base to achieve safe migration of decision-making strategies in different penetration scenarios;

[0108] The verification layer is used to build a digital twin verification platform and implement strategy rehearsal and effect tracing.

[0109] A computer-readable storage medium stores a computer program, which, when executed, implements the method for controlling voltage of a distribution network under uncertain factors.

[0110] To validate the performance of this invention, this example simulates a mixed industrial and residential load scenario with a 35% photovoltaic penetration rate based on a 33-node IEEE distribution network system. The data source uses 5-minute real-world operational data, including sudden weather changes such as cloudy skies and thunderstorms. The system baseline capacity is 12.5 MVA, and the line power threshold is set at 88%. The XGBoost model is used for comparison.

[0111] Figure 2 The effects of distribution network operation situation awareness under source-load fluctuation scenarios were compared:

[0112] 1) At 10:48, the photovoltaic output suddenly changed due to the rapid movement of clouds (red star), and the power fluctuated by ±9% within 20 minutes;

[0113] 2) At 19:12, the start and stop of the industrial production line caused a surge in load (purple cross mark), resulting in an 8% step change.

[0114] The root mean square error of the curve of the present invention (red dotted line) at the mutation point is 1.8%, which is 57% lower than that of the XGBoost method (blue dotted line). In addition, the situational awareness stability during periods of drastic power gradient changes is significantly better, laying a foundation for accurate decision-making information for the distribution network auxiliary decision-making system.

[0115] Figure 3 The effect of distribution network decision-making assistance is demonstrated. This figure shows that the present invention is significantly superior to traditional methods and manual scheduling in terms of network loss control, safety risk prevention and control, and new energy consumption. By integrating the source-load fluctuation correlation characteristics with multi-objective collaborative optimization, the system can still maintain efficient decision-making in extreme scenarios such as typhoons, effectively balancing economy and safety. Compared with the limitations of traditional single-objective optimization and manual experience, the present invention has constructed an economic, safe, and green collaborative distribution network decision-making assistance system, verifying the strong adaptability and comprehensive benefits of the dynamic strategy system.

Claims

1. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations, characterized in that: The following steps are involved: S1: Collect distribution network substation information, quantify the correlation between the synchronization and hysteresis of multi-source heterogeneous data in PV output and load power fluctuations, calculate the composite feature vector based on this correlation, and establish a standardized risk perception dataset containing aggregated information of distribution network substations; S2: Define the state space and action space of the reinforcement learning algorithm based on the composite feature vector, design the reward function, and build a deep reinforcement learning model architecture to assist in decision-making and optimization of the distribution network. S3: Build a scene feature library. For a new target scene, calculate its fluctuation correlation similarity with historical scenes in the feature library, select the reference scene set with the strongest correlation, use the learning algorithm to train the deep reinforcement learning model architecture, freeze the feature extraction layer parameters, and retain its ability to analyze source-load correlation features. Use a composite reward function to constrain the transfer training direction to achieve rapid scene adaptation. S4: Build a digital twin simulation platform for the power grid and design evaluation indicators in different dimensions. Input the real-time distribution network substation dataset into the migration-trained deep reinforcement learning model architecture to generate several groups of differentiated candidate solutions. Use the digital twin simulation platform to deduce all solutions and compare some of them to obtain recommended strategies. After successful verification, the recommended strategies will be stored in the strategy knowledge base. When there are similar scenarios, the strategies will be executed first.

2. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 1, characterized in that: The step S1 includes the following sub-steps: S101. Multi-source heterogeneous data collection and cleaning: Collect distribution network area related information, including but not limited to distribution network area aggregated information, distributed power output time series data, load power curves, meteorological information and network topology parameters, historical over-limit event records and environmental variables of the corresponding time period, and use the quartile method to eliminate outliers in the distribution network area related information. Then, use the sliding window mean interpolation and adjacent time period similarity matching algorithm to repair the missing data. S102. Data alignment and feature decoupling in spatiotemporal dimensions: Interpolation alignment is performed on the collected multi-source heterogeneous data at a unified time resolution. Geographic information matching technology is used to correlate the load data of distributed power generation nodes and power supply areas. Subsequently, the time series data is segmented using a sliding time window to extract statistical features of source-load fluctuations within each time window, including but not limited to extreme values, variance, and volatility. Subsequently, the fluctuation components of weather-sensitive loads and rigid loads are separated based on the correlation between weather changes and load changes. S103. Quantitative Modeling of Fluctuation Correlation Characteristics: Grey correlation analysis is used to calculate the correlation index between source-side output fluctuations and load-side power changes at different time scales. A fluctuation pattern matching library is further constructed based on the correlation index to identify the coupling characteristics of source-load fluctuations in the time-frequency domain. A joint probability distribution model of source-load fluctuations is established using a non-parametric kernel density estimation method to quantify the correlation between the synchronization and hysteresis of the two fluctuations. This joint probability distribution model of source-load fluctuations is established using a non-parametric kernel density estimation method. S104. Composite feature vector construction and label definition: The synchronization coefficient and hysteresis coefficient of the fluctuation characteristics of the original data including photovoltaic output and load power are integrated in multiple dimensions to obtain a composite feature vector including timestamp, spatial location, fluctuation pattern, and correlation strength. The risk level label is then defined based on the deviation between the actual power of the line and the capacity threshold. Dynamic threshold correction technology is used to eliminate the influence of seasonal characteristics on the label distribution. S105. Scenario-based dataset division and enhancement: The dataset is classified into scenarios according to meteorological type, load characteristics, and network topology. Then, for the scarce risk scenario samples and small sample scenarios after classification, the SMOTE oversampling algorithm and the virtual sample generation strategy based on the adversarial generative network are used to perform data enhancement. Finally, the validity of the associated features is verified to obtain a standardized risk perception dataset with strong representation capabilities.

3. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 1, characterized in that: In step S103, the synchronicity of the fluctuations is expressed as: Where I(X,Y) is the synchronization coefficient of the discretized observations of the source-side power fluctuation sequence X and the load-side power fluctuation sequence Y; p(x,y) is the joint probability distribution of element x in X and element y in Y; p(x) and p(y) are the marginal probability distributions of X and Y. The hysteresis of the fluctuation is expressed as: Among them, τ is the time lag parameter; R XY is the hysteresis coefficient of the source side power fluctuation sequence X and the load side power fluctuation sequence Y; t ,Y t+τ are the fluctuation values ​​of the source side at time t and the fluctuation values ​​of the load side at time t+τ respectively; μ X ,μ Y are the means of the source-load fluctuation series respectively; t is the time value; N is the total number of data points in the source-load side power fluctuation time series.

4. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 2, characterized in that: In step S105, when verifying the validity of the associated features: a feature contribution analysis method based on Shapley value is used to quantify the risk perception contribution weight of each associated feature to the over-limit risk, and by comparing the difference in model risk perception performance of the actual power grid operation data with and without the associated features, redundant features are eliminated and the feature combination is optimized.

5. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 1, characterized in that: The step S2 includes the following sub-steps: S201, Dynamic State Space Construction and Feature Encoding: Based on the composite feature vector generated in step S1, the state space of the reinforcement learning algorithm is defined, and the state vectors added include but are not limited to: distribution network area aggregation information, real-time source-load fluctuation correlation indicators, historical line crossing frequency, current environmental variables, and topological structure characteristics; S202, Action Space Design and Risk Level Mapping: The action space of the reinforcement learning algorithm is defined by pre-delineating the probability intervals for network operation risk perception. The intervals cover low, medium, and high risk levels, each corresponding to a warning level under a specific probability threshold. An adaptive action classification mechanism is introduced to dynamically adjust the interval boundaries based on the line load factor. S203, Deep Reinforcement Learning Network Architecture Implementation: Construct a hybrid neural network based on the LSTM-AC framework. The executor network uses a bidirectional LSTM structure to capture the temporal correlation characteristics of source-load fluctuations and output action probability distributions. The evaluator network integrates a graph attention mechanism to analyze the spatial constraints of the distribution network topology on the line flow. S204. Experience replay and priority sampling training strategy: Deploy a dynamic experience pool to store interaction data, and use a priority sampling mechanism to give a higher selection probability to training samples in high-risk periods.

6. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 5, characterized in that: In the reinforcement learning described above, a reward function integrating multi-objective optimization is designed, which can be specifically expressed as: Among them, R t The total reward value at time t in the direction of guiding strategy optimization; is an indicator function, when the risk perception state a t and the actual limit-crossing state y t If they are consistent, it is 1, otherwise it is 0; ω1, ω2, ω3 are weight coefficients; P t pred is the predicted line power at time t; P t actual is the actual line power at time t; P th is the line safety power threshold; K is the look-ahead time step.

7. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 5, characterized in that: In step S204, a phased training strategy is designed: in the early stage, exploratory training is emphasized, and the action coverage is broadened by injecting Gaussian noise; in the middle and late stages, the utilization training weight is gradually increased, and the Monte Carlo tree search optimization strategy is used to select the path. Adversarial sample training is implemented simultaneously to improve the model's robustness to abnormal fluctuations.

8. The distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 1 is characterized in that: The step S3 includes the following sub-steps: S301. Scenario feature library construction: Based on the standardized data set established in step S2, source-load fluctuation correlation characteristics and environmental parameters are extracted for various scenarios. Scenarios include, but are not limited to, high photovoltaic power generation and sudden load changes. Source-load fluctuation correlation characteristics include, but are not limited to, synchronization coefficients and lagged correlation coefficients. Based on these data, a multi-dimensional scenario feature library is constructed, including fluctuation patterns, network topology, and meteorological conditions, as a basic carrier of transferable knowledge. S302. Target scenario feature similarity assessment: For the new target scenario, calculate its fluctuation correlation similarity with the historical scenarios in the feature library in S301, quantify the consistency of the source-load fluctuation pattern, and simultaneously select the reference scenario set with the strongest correlation as the data basis for migration; S303, scene migration: Reuse the deep reinforcement learning model architecture trained in step S2, freeze the feature extraction layer parameters, retain its ability to analyze source-load correlation features, fine-tune only the fully connected layer of the policy network for a small amount of labeled data of the target scene, and use the composite reward function designed in step S2 to constrain the migration training direction to achieve rapid scene adaptation.

9. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 8, characterized in that: The similarity evaluation model between the target new scene and the source scene is specifically as follows: Among them, D KL is the distribution difference measure between the source scene and the target scene; P(x) is the joint probability distribution of the source scene; Q(x) is the joint probability distribution of the target scene, that is, the estimated value of the new scene; The feature space includes but is not limited to source-load correlation indicators and fluctuation patterns.

10. A distribution network auxiliary decision-making method considering the correlation between source and load fluctuations according to claim 1, characterized in that: The step S4 includes the following sub-steps: S401. Construction of a simulation environment for auxiliary decision-making strategies: Based on the source-load fluctuation correlation dataset established in step S1 and the auxiliary decision-making model trained in step S2, a power grid digital twin simulation platform is constructed to support dynamic deduction of auxiliary decision-making strategies. S402. Design of a multi-dimensional evaluation index system: Design evaluation indicators for economy, safety, and renewable energy absorption rate. Economy includes, but is not limited to, network losses and dispatch instruction execution costs; safety includes, but is not limited to, over-limit risk attenuation rate and voltage compliance rate; and renewable energy utilization includes, but is not limited to, curtailed solar power rate, curtailed wind power rate, and load matching. S403, intelligent generation of candidate scheduling strategies: In combination with the migration deep reinforcement learning model architecture model of step S3, a candidate strategy set is automatically generated based on the current over-limit risk level, including but not limited to: energy storage charging and discharging power optimization, interruptible load regulation, and distributed power output control, and then multiple sets of differentiated candidate solutions are generated; S404. Dynamic simulation verification of strategy effectiveness: Utilize the power grid digital twin simulation platform established in step S401 to calculate all differentiated candidate solutions generated in step S403, and collect the evaluation indicators of the multi-dimensional evaluation indicator system in step S402 as key indicator time series data, as well as the over-limit risk value, new energy consumption curve smoothness, and node voltage fluctuation rate generated within one hour after the solution is implemented. These are then compared with the undispatched baseline scenario and scored; S405. Comprehensive comparison and optimization of multiple strategies: Evaluate the long-term benefits of each strategy based on the subsequent risk evolution trend of the decision-making support to set the strategy validity period attenuation factor. Select the three most consistent strategies with significant short-term benefits and controllable long-term negative impacts from the multiple sets of differentiated candidate solutions generated in S403 as recommended strategies and generate a quantitative benefit comparison table. S406. Dynamic evaluation report generation and feedback: Based on the recommended strategy and the quantitative benefit comparison table from S405, an evaluation report containing a strategy execution plan, a map of expected effects, and risk warning information is generated and presented to the operator. The operator then manually verifies the report and waits for the verification results of the recommended strategy. S407, iterative update of the strategy knowledge base: the verified effective recommended strategies are classified and stored in the strategy knowledge base according to the scenario characteristics. When a similar scenario triggers an out-of-limit warning, the strategy is executed first or the parameters are fine-tuned based on the real-time effect.

11. A distribution network auxiliary decision system considering the correlation between source and load fluctuations, comprising a perception layer, a decision layer, a migration layer, and a verification layer, characterized in that: The perception layer integrates distribution network area aggregation information and area-level edge computing nodes to collect real-time output fluctuations of distributed power sources, distribution transformer load rates, and user-side response potential; The decision-making layer is deployed with a cloud-edge collaborative reinforcement learning decision engine to generate an optimization strategy that considers the spatiotemporal correlation between source and load; The migration layer is used to establish a scene feature knowledge base to achieve safe migration of decision strategies in different penetration scenarios; The verification layer is used to build a digital twin verification platform to implement strategy preview and effect tracing.

12. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed, a method for controlling voltage of a distribution network under uncertain factors according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Virtual power plant optimal scheduling method considering renewable energy sources

    CN120810609A

  • A virtual power plant optimal scheduling method considering renewable energy

    CN120810609B

  • Business auditing method and system based on large model, electronic equipment and medium

    CN120875816A

  • Multi-modal large model-based intelligent operation and maintenance method and system for data medium station

    CN120912010A

  • Power distribution network planning method based on decision-oriented learning

    CN121031996A