Multi-objective optimization scheduling method and system for virtual power plants integrated with deep reinforcement learning
Through deep reinforcement learning agents to calculate the voltage sensitivity and weak lines of grid nodes in virtual power plants, they generate dynamic allocation strategies for reactive resource, which solves the real-time response and adaptive adjustment problems of virtual power plants under the topology of the grid, and improves the voltage stability and economy of the power grid.
Patent Information
- Application Number
- CN202510858341.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the face of frequent changes in the grid topology and high-dimensional state space after distributed power supply access, existing virtual power plant scheduling methods are difficult to achieve real-time response and online adaptive adjustment, resulting in insufficient model mismatch and control accuracy, and it is difficult to meet the real-time and coordination requirements of voltage stability control and dynamic allocation of reactive resource.
Deep reinforcement learning agents are adopted to obtain short-circuit capacity data of network connection points and real-time impedance spectrum measurement data of virtual power plants, calculate the voltage sensitivity parameters of power grid nodes, locate weak line segments, and generate dynamic allocation strategies for reactive resource, with the goal of maximizing the profit of voltage regulation compensation and minimizing equipment regulation loss, and coordinate the power of distributed power and energy storage converters.
Real-time response and adaptive adjustment in complex and variable power grid environments are realized, the grid voltage stability and operational economy are improved, the virtual power plants have the need for fast adaptive scheduling, and the low search efficiency and model mismatch problems of traditional methods in high-dimensional and nonlinear environments are overcome.
Smart Images

Figure CN120377361B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual power plant scheduling, and in particular to a multi-objective optimization scheduling method and system for a virtual power plant integrating deep reinforcement learning. Background Art
[0002] With the ongoing transformation of the energy mix and the continued increase in renewable energy penetration, virtual power plants (VPPs), as a key technological means of integrating distributed energy resources and enhancing grid operational flexibility, are gradually playing a critical role in power systems. In practical applications, VPPs must implement multi-objective optimized scheduling within complex and volatile grid environments, facing significant challenges in voltage stability control and the dynamic allocation of reactive power resources. Due to the dispersed access locations and high volatility of distributed power sources, traditional centralized scheduling methods struggle to meet real-time and collaborative requirements.
[0003] A typical approach currently relies on a combination of model predictive control and heuristic algorithms. This approach constructs mathematical models of each unit within a virtual power plant and uses heuristic algorithms such as genetic algorithms or particle swarm optimization during a rolling optimization process to find the optimal reactive power allocation solution. However, existing solutions still have significant limitations. For example, they rely on precise system modeling, but the frequent changes in grid topology after the integration of distributed energy resources lead to significant model mismatch problems, which affect control accuracy. Heuristic algorithms also have low search efficiency in high-dimensional state spaces, making real-time response and online adaptive adjustment difficult. Summary of the Invention
[0004] The present invention provides a virtual power plant multi-objective optimization scheduling method and system integrating deep reinforcement learning, which is used to solve the problems of prominent model mismatch and insufficient scheduling accuracy in the prior art, as well as the difficulty in achieving real-time response and online adaptive adjustment.
[0005] In a first aspect, the present invention provides a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning, comprising:
[0006] Obtain short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point;
[0007] Calculating voltage sensitivity parameters of each grid node in the access area of the virtual power plant based on the short-circuit capacity data;
[0008] locating a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data;
[0009] Based on the voltage sensitivity parameter and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize voltage regulation compensation benefits and minimize equipment regulation losses;
[0010] According to the reactive resource dynamic allocation strategy, the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter are coordinated and controlled to achieve multi-objective optimization scheduling of the virtual power plant.
[0011] Optionally, calculating the voltage sensitivity parameter of each grid node in the access area of the virtual power plant based on the short-circuit capacity data includes:
[0012] Calculate the cumulative value of the line impedance from the virtual power plant's grid connection point to each grid node in the virtual power plant's access area;
[0013] Extracting a target impedance component whose proportion exceeds a set threshold value from the accumulated value of the line impedance, and using the target impedance component as a dominant impedance component;
[0014] Calculating a basic voltage sensitivity factor of each grid node in the access area of the virtual power plant based on the dominant impedance component and the short-circuit capacity data of the grid connection point;
[0015] Calculating a weighting factor based on the real-time current carrying capacity and the rated current carrying capacity parameters of the line associated with the dominant impedance component;
[0016] The basic voltage sensitivity factor is superimposed on the weight factor to generate the voltage sensitivity parameter of each grid node in the access area of the virtual power plant.
[0017] Optionally, calculating the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant according to the dominant impedance component and the short-circuit capacity data of the grid connection point includes:
[0018] Obtaining conductor material property parameters, ambient temperature measured value, spatial length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component;
[0019] Calculating a reference resistivity value and a temperature drift compensation coefficient based on the conductor material property parameters and the measured ambient temperature value;
[0020] Calculating line geometric characteristic values according to the spatial length dimension and conductor cross-sectional area parameters;
[0021] Performing a product operation on the line geometric characteristic value and the temperature drift compensation coefficient to obtain a physical correction factor;
[0022] The electrical characteristic factor is calculated based on the dominant impedance component and the short-circuit capacity data of the grid connection point, and combined with the physical correction factor, the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant is calculated.
[0023] Optionally, calculating a weighting factor according to the real-time current carrying capacity and rated current carrying parameters of the line associated with the dominant impedance component includes:
[0024] Obtaining the real-time current carrying capacity, rated current carrying capacity parameters and historical maximum current carrying capacity records of the line associated with the dominant impedance component;
[0025] Calculating the instantaneous approximation ratio of the real-time current carrying capacity to the rated current carrying capacity parameter, and determining the current carrying capacity fluctuation coefficient based on the historical maximum current carrying capacity record;
[0026] Obtaining the conductor surface contamination level and environmental temperature and humidity parameters of the line associated with the dominant impedance component;
[0027] According to the contamination level of the conductor surface, determine the contamination thermal resistance increment coefficient, and calculate the air heat dissipation efficiency in combination with the ambient temperature and humidity parameters;
[0028] Calculating a dominant impedance reference value based on the dominant impedance component, and obtaining a high-frequency harmonic distortion rate of the line associated with the dominant impedance component to calculate a skin effect depth correction coefficient;
[0029] The instantaneous approach ratio, the current fluctuation coefficient, the pollution thermal resistance increment coefficient, the air heat dissipation efficiency and the skin effect depth correction coefficient are multiplied to obtain a comprehensive current carrying capacity factor, and the weight factor is calculated in combination with the dominant impedance reference value.
[0030] Optionally, locating a weak line section associated with a power grid node according to the real-time impedance spectrum measurement data includes:
[0031] Analyzing the real-time impedance spectrum measurement data to obtain impedance amplitude and phase angle change;
[0032] Calculating the fluctuation coefficient of the phase angle variation within a preset time window, and simultaneously calculating the deviation of the impedance amplitude from a historical reference amplitude, and marking the corresponding line segment where the deviation exceeds an impedance anomaly threshold and the fluctuation coefficient exceeds a preset stability threshold as a candidate weak segment;
[0033] encoding the candidate weak section into a directed line segment according to the starting grid node and the ending grid node of the candidate weak section;
[0034] Traverse all directed line segments. If the end node of the first directed line segment is the same as the start node of the second directed line segment, merge the first directed line segment and the second directed line segment into a target line segment. Combine all target line segments to generate a weak line segment.
[0035] Optionally, based on the voltage sensitivity parameter and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, including:
[0036] Creating a feature group for each grid node, the feature group comprising a node optimization sensitivity parameter and a weak node, the weak node being determined based on whether the corresponding node is a starting grid node or an ending grid node of a weak line segment;
[0037] Determine the target grid node to which each reactive resource device is connected according to a preset device node association table;
[0038] Utilizing a pre-trained deep reinforcement learning agent, the feature group and the target grid node are analyzed to generate a regulation instruction value for each reactive resource device;
[0039] The regulation command value is converted into a reactive power setting value of the distributed power source and a charge and discharge power setting value of the energy storage converter, so as to generate a reactive resource dynamic allocation strategy according to the reactive power setting value and the charge and discharge power setting value.
[0040] Optionally, a feature group is created for each grid node, the feature group including a node optimization sensitivity parameter and a weak node, the weak node being determined based on whether the corresponding node is a starting grid node or an ending grid node of a weak line section, including:
[0041] Collect voltage fluctuation trajectory data of grid nodes within a preset time window to calculate the voltage offset integral value;
[0042] Calculate the connection weight coefficients of the starting grid node and the ending grid node based on the node connection relationship in the weak line section;
[0043] The connection weight coefficient corresponding to the grid node marked as a weak node is multiplied by the voltage offset integral value to generate a risk enhancement factor, and the node optimization sensitivity parameter is calculated in combination with the voltage sensitivity parameter;
[0044] The node optimization sensitivity parameters and the weak nodes are combined into a feature group core element, and the feature group core element is combined with the hierarchical position code of the power grid node in the virtual power plant to generate a feature group for each power grid node.
[0045] In a second aspect, the present invention provides a virtual power plant multi-objective optimization scheduling system integrating deep reinforcement learning, comprising:
[0046] An acquisition module is used to obtain short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point;
[0047] a calculation module, configured to calculate a voltage sensitivity parameter of each grid node in an access area of the virtual power plant based on the short-circuit capacity data;
[0048] a positioning module, configured to locate a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data;
[0049] A generation module is used to generate a dynamic reactive resource allocation strategy based on the voltage sensitivity parameters and the information of the weak line section using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the equipment regulation loss.
[0050] The control module is used to coordinately control the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter according to the dynamic allocation strategy of reactive resources, so as to achieve multi-objective optimization scheduling of the virtual power plant.
[0051] In a third aspect, an embodiment of the present invention provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning as described in the first aspect above.
[0052] In a fourth aspect, an embodiment of the present invention provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning as described in the first aspect.
[0053] In the present invention, short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point are obtained; based on the short-circuit capacity data, the voltage sensitivity parameters of each grid node in the access area of the virtual power plant are calculated; according to the real-time impedance spectrum measurement data, the weak line section associated with the grid node is located; according to the voltage sensitivity parameters and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, and the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the equipment regulation loss; according to the dynamic reactive resource allocation strategy, the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter are collaboratively controlled to achieve multi-objective optimal scheduling of the virtual power plant. The technical solution provided by the present invention breaks through the traditional method's reliance on fixed power grid models, perceives the dynamic topological changes and impedance characteristics of the power grid in real time, provides a dynamic data basis for subsequent calculations to adapt to the fluctuations of renewable energy, and solves the model mismatch problem caused by the access of distributed power sources; quantifies the sensitivity of the voltage of each node in the power grid to the change of reactive power, accurately identifies the key nodes with high voltage regulation efficiency, provides a scientific basis for the optimal allocation of resources, and overcomes the defects of traditional centralized scheduling that are difficult to deal with the dispersion and volatility of distributed power sources; actively identifies the vulnerable links (i.e., weak line sections) in the power grid that are prone to voltage instability or power quality problems, focuses the optimization focus on the risk areas, and improves the overall operation of the system. The intelligent agent can improve the robustness of the network and make up for the shortcomings of traditional methods in not being able to perceive local power grid risks. Without the need for precise mathematical models, the intelligent agent can interact with the environment and learn online to efficiently coordinate the dual goals of maximizing the voltage regulation compensation benefit and minimizing the equipment regulation loss. This can solve the core problems of heuristic algorithms in high-dimensional, nonlinear, and highly uncertain environments, such as low search efficiency, difficulty in real-time response, and adaptive adjustment. The intelligent agent can transform the optimization strategy into direct and coordinated control instructions for the reactive power of distributed power sources and energy storage converters, so as to realize the rapid aggregation response and coordinated optimization scheduling of distributed resources in the virtual power plant, improve the voltage stability and operation economy of the power grid, and meet the real-time and coordination requirements in a complex and changeable power grid environment.Furthermore, the voltage offset integral value is calculated based on the voltage fluctuation trajectory of the grid node, and the risk enhancement factor is calculated in combination with the connection weight of the node in the weak line section; secondly, this factor is integrated with the voltage sensitivity parameter to generate a node optimization sensitivity parameter that can better reflect the optimization value and risk; then, the node optimization sensitivity parameter, the weak node status and the hierarchical position encoding of the node in the virtual power plant are combined to form a feature group for each node; at the same time, according to the preset device node association table, the target grid node to which each reactive resource device is connected is determined; finally, the pre-trained deep reinforcement learning agent is used to analyze these feature groups and target grid nodes, and directly output The optimal regulation command value for each reactive resource device is then converted into a specific power setting value to form a dynamic allocation strategy, enabling the reinforcement learning agent to deeply understand the dynamic characteristics and regulation priorities of the power grid in an environment where model parameters are unknown or changing, overcoming the problem of decreased control accuracy caused by model mismatch and topology changes in traditional methods; converting high-dimensional, complex multi-objective optimization problems into an agent decision-making process based on rich features, avoiding the inefficient search of heuristic algorithms in high-dimensional space, greatly improving the real-time performance and computational efficiency of generating high-quality, executable strategies in a high-uncertainty environment, and meeting the urgent needs of virtual power plants for fast adaptive scheduling.
[0054] These and other aspects of the present invention will become more readily apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 A flowchart of a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning provided by the present invention is shown;
[0057] Figure 2 A schematic diagram of the structure of a virtual power plant multi-objective optimization scheduling system integrating deep reinforcement learning provided by the present invention is shown;
[0058] Figure 3 A schematic structural diagram of a computing device provided by the present invention is shown. DETAILED DESCRIPTION
[0059] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0060] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0062] In view of the problems that the existing virtual power plant optimization scheduling methods based on model predictive control and heuristic algorithms face, such as frequent changes in grid topology, decreased system modeling accuracy, and low search efficiency in high-dimensional state space, it is difficult to meet the real-time and collaborative requirements of voltage stability control and dynamic allocation of reactive resources in complex and changing environments. This paper proposes a virtual power plant multi-objective optimization scheduling method that integrates deep reinforcement learning. This method makes scheduling no longer rely on precise mathematical modeling. Instead, it introduces a deep reinforcement learning agent and combines the short-circuit capacity and real-time impedance spectrum measurement data of the virtual power plant grid connection point to extract key voltage sensitivity parameters and identify weak line sections, thereby constructing an adaptive decision-making mechanism guided by maximizing voltage regulation compensation benefits and minimizing equipment regulation losses. By online learning of grid operating status and environmental changes, the deep reinforcement learning agent can achieve coordinated control of distributed power source reactive output and energy storage converter charging and discharging power, improve the response speed and multi-objective optimization capability of the scheduling strategy, overcome the limitations of traditional methods in dynamic adaptability and computational efficiency, and meet the urgent needs of virtual power plants in new power systems for intelligent and real-time scheduling.
[0063] Figure 1 The present invention provides a flowchart of a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning, such as Figure 1 As shown, the method includes:
[0064] Step 101: Acquire short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point;
[0065] In this step, the virtual power plant refers to a cloud-based control system that aggregates distributed power sources, energy storage systems, and controllable loads. It includes photovoltaic inverters, energy storage converters, and load controllers, achieving wide-area coordinated control via a communication network. The grid connection point refers to the physical connection point between the virtual power plant and the public grid, typically a 10kV / 35kV substation busbar, equipped with voltage and current transformers for electrical parameter measurement. Short-circuit capacity data refers to the apparent power value (MVA) when a three-phase short circuit occurs at the grid connection point. This data is calculated using a comprehensive power system analysis program or power system analysis software package and reflects the inverse of the grid's equivalent impedance. Real-time impedance spectrum measurement data refers to the impedance amplitude (Ω) and phase angle (°) in the 0.1-10kHz frequency band. This data is obtained using a frequency sweep method and contains data at 512 frequency points.
[0066] In an embodiment of the present invention, the short-circuit capacity data of the virtual power plant's grid connection point is collected by the data acquisition and monitoring system of the power grid dispatching center. This data comes from the power system short-circuit calculation database. At the same time, a distributed frequency response analyzer is used to perform sweep frequency excitation (0.1-10kHz) on the grid connection point and obtain real-time impedance spectrum measurement data.
[0067] Step 102: Calculating voltage sensitivity parameters of each grid node in the access area of the virtual power plant based on the short-circuit capacity data;
[0068] In this step, the access area refers to the distribution network managed by the virtual power plant. The topology is either radial or ring, consisting of 35 grid nodes and 28 feeders. A grid node is a uniquely numbered line connection point or device access point in the distribution network. It is equipped with a smart meter for electrical quantity monitoring. The voltage sensitivity parameter is the ratio of the node voltage change rate to the change in reactive power injection (unit: V / kVar), with a value range of 0.02-0.15.
[0069] In an embodiment of the present invention, the numerical value of the short-circuit capacity data is first analyzed, and the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant is calculated based on the target impedance component of the accumulated value of the line impedance from the grid connection point of the virtual power plant to each grid node in the access area of the virtual power plant, which accounts for more than a set threshold; based on the real-time current carrying capacity and rated current carrying capacity parameters of the line associated with the dominant impedance component, a weight factor is calculated, and the weight factor is superimposed on the basic voltage sensitivity factor to generate the voltage sensitivity parameter of each grid node in the access area of the virtual power plant.
[0070] Step 103: locating a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data;
[0071] In this step, the weak line section refers to the physical line section where the impedance abnormally increases by more than 20%, which is defined by the coordinates of the starting and ending grid nodes and is usually 0.5-3 km in length.
[0072] In this embodiment of the present invention, a fast Fourier transform is performed on real-time impedance spectrum measurement data to extract the fundamental impedance amplitude and phase angle variation. The fluctuation coefficient of the phase angle variation within a preset time window is calculated using a sliding time window. The deviation of the impedance amplitude from a historical baseline amplitude is compared. Line segments where the deviation exceeds an impedance anomaly threshold and the fluctuation coefficient exceeds a preset stability threshold are marked as candidate weak segments. Based on a graph theory algorithm, target line segments within all candidate weak segments are merged based on node connectivity to generate a weak line segment.
[0073] Step 104: Based on the voltage sensitivity parameter and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize voltage regulation compensation benefits and minimize equipment regulation losses;
[0074] In this step, the reactive resource dynamic allocation strategy refers to the set of control instructions output by the deep reinforcement learning agent, and the data structure is in JSON format.
[0075] In an embodiment of the present invention, a feature group is created for each grid node, and the target grid node to which each reactive resource device is connected is determined based on a preset device node association table; a pre-trained deep reinforcement learning agent is used to analyze the feature group and the target grid node, generate an adjustment instruction value for each reactive resource device, and convert it into a reactive power setting value of a distributed power source and a charge and discharge power setting value of an energy storage converter, and generate a dynamic reactive resource allocation strategy based on the reactive power setting value and the charge and discharge power setting value.
[0076] Step 105: According to the reactive resource dynamic allocation strategy, the reactive output power of the distributed power source and the charge and discharge power of the energy storage converter are collaboratively controlled to achieve multi-objective optimization scheduling of the virtual power plant;
[0077] In this step, reactive output power refers to the capacitive or inductive reactive power provided by the distributed generation (kVar). The PV inverter's regulation range is -100% to +100% of the rated capacity. Charge and discharge power refers to the active power command of the energy storage converter (kW). Positive values indicate discharge, negative values indicate charge, and the limit is ±1C.
[0078] In this embodiment of the present invention, the set value instructions in the reactive resource dynamic allocation strategy are parsed, and the reactive power set value (unit: kVar) is sent to the photovoltaic inverter through the substation communication network and system. The charging and discharging power instructions (unit: kW) are sent to the energy storage converter via Modbus TCP. Specifically, a master-slave control architecture is adopted to ensure the timing synchronization of the power instructions of multiple devices, and the voltage at the grid connection point of the virtual power plant is stabilized in the range of 0.95-1.05pu.
[0079] The embodiments of the present invention solve the complex problems of inaccurate regulation resource allocation, frequent equipment operation and insufficient economy in traditional methods, and achieve the comprehensive optimization goals of voltage qualification rate and equipment loss.
[0080] The present invention provides a specific embodiment, step 102, calculating the voltage sensitivity parameter of each grid node in the access area of the virtual power plant based on the short-circuit capacity data, specifically includes the following steps:
[0081] Step 201: Calculate the cumulative value of the line impedance from the virtual power plant's grid connection point to each grid node in the virtual power plant's access area;
[0082] In this step, the line impedance accumulation value refers to the arithmetic sum (unit: pu) of all line impedance moduli on the path from the grid connection point to the target grid node. It is calculated and accumulated based on the line resistance and reactance parameters, and reflects the electrical distance.
[0083] In this embodiment of the present invention, a topological path tracing algorithm is used to traverse the distribution network step by step, starting from the virtual power plant's grid connection point to the target grid node. The impedance values (unit: Ω) of all line segments along the path are accumulated to obtain the accumulated line impedance value. The specific process involves accessing the power grid geographic information system database to obtain line resistance and reactance parameters, and then performing arithmetic accumulation of the impedance modulus values of each line along the path.
[0084] Step 202: extracting a target impedance component whose proportion exceeds a set threshold value from the accumulated line impedance value, and taking the target impedance component as a dominant impedance component;
[0085] In this step, the threshold is a preset impedance contribution threshold (typically 15%), which is used to screen critical paths. The target impedance component is the impedance value of a single path within the path. Its proportion of the cumulative value must exceed the set threshold, indicating a significant contribution to the total impedance. The dominant impedance component is the vector sum of the target impedance components (unit: Ω) and represents the equivalent impedance of the critical path.
[0086] In this embodiment of the present invention, a contribution analysis method is used to process the accumulated impedance value. Specifically, the impedance of each line is divided by the accumulated line impedance value to obtain a contribution. Assuming a threshold (typically 15%), line impedance components exceeding this threshold are screened out, such as line L8, whose impedance contribution is 18%. These components are extracted to form the target impedance component, and their vector sum is taken to generate the dominant impedance component (unit: Ω).
[0087] Step 203: Calculating the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant based on the dominant impedance component and the short-circuit capacity data of the grid connection point;
[0088] In this step, the short-circuit capacity data refers to the apparent power (in MVA) at the grid connection point during a three-phase short circuit, calculated using power system transient stability analysis software. The basic voltage sensitivity factor refers to the preliminarily calculated voltage-reactive power sensitivity benchmark (dimensionless).
[0089] In an embodiment of the present invention, the conductor material property parameters, the measured value of the ambient temperature, the spatial length dimension, and the conductor cross-sectional area parameters of the line corresponding to the dominant impedance component are obtained to calculate the reference resistivity value, the temperature drift compensation coefficient, and the line geometric characteristic value; the line geometric characteristic value and the temperature drift compensation coefficient are multiplied to obtain a physical correction factor; the electrical characteristic factor is calculated based on the short-circuit capacity data of the dominant impedance component and the grid connection point, and combined with the physical correction factor, the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant is calculated.
[0090] Step 204: Calculate a weight factor based on the real-time current carrying capacity and rated current carrying parameters of the line associated with the dominant impedance component;
[0091] In this step, the real-time current capacity refers to the effective current value (A) measured by the line current transformer. It is updated every 5 seconds and reflects the instantaneous load status of the line. The rated current capacity parameter refers to the maximum continuous current capacity (A) allowed by the line design. It is determined during the grid planning phase based on the conductor material and heat dissipation conditions. The weighting factor is a coefficient (0.5-1.5) that corrects the basic sensitivity. It is calculated by combining the real-time current ratio, the incremental coefficient of the contamination thermal resistance, the heat dissipation efficiency, and the harmonic effects. It is used to quantify the line safety margin.
[0092] In an embodiment of the present invention, the real-time current carrying capacity, rated current carrying parameters, and historical maximum current carrying records of a line associated with a dominant impedance component are obtained, the instantaneous approximation ratio of the real-time current carrying capacity to the rated current carrying parameters is calculated, and the current fluctuation coefficient is determined based on the historical maximum current carrying record; the surface contamination level of the conductor and the ambient temperature and humidity parameters of the line associated with the dominant impedance component are obtained to calculate the contamination thermal resistance increment coefficient and the air heat dissipation efficiency; the skin effect depth correction coefficient is calculated based on the dominant impedance component; the instantaneous approximation ratio, the current fluctuation coefficient, the contamination thermal resistance increment coefficient, the air heat dissipation efficiency, and the skin effect depth correction coefficient are multiplied to obtain a comprehensive current carrying capacity factor, and the weight factor is calculated in combination with the dominant impedance reference value.
[0093] Step 205: superimposing the basic voltage sensitivity factor and the weight factor to generate a voltage sensitivity parameter of each grid node in the access area of the virtual power plant;
[0094] In the embodiment of the present invention, the basic voltage sensitivity factor and the weight factor are superimposed and calculated according to a linear weighted formula, and the result is normalized to a standard interval [0, 1], so as to finally obtain the voltage sensitivity parameter.
[0095] The embodiment of the present invention solves the problem of poor adaptability of traditional centralized scheduling methods to changes in grid topology. By extracting the dominant impedance component and dynamically correcting the voltage sensitivity in combination with the weight factor, it can more accurately reflect the voltage sensitivity characteristics of the grid nodes after the distributed power source is connected, thereby improving the voltage stability control capability of the virtual power plant in a complex grid environment and reducing the control deviation caused by model mismatch.
[0096] The present invention provides a specific embodiment, step 203, calculating the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant based on the dominant impedance component and the short-circuit capacity data of the grid connection point, specifically comprising the following steps:
[0097] Step 211: Obtaining conductor material property parameters, ambient temperature measured value, space length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component;
[0098] In this step, conductor material attribute parameters refer to the coded data describing the physical characteristics of the line conductor, including material type, alloy code, and resistance temperature coefficient. The measured ambient temperature refers to the real-time air temperature (unit: °C) at the line location. This value is collected by a platinum resistance temperature sensor and reflects the heat dissipation conditions of the conductor. The spatial length dimension refers to the straight-line distance between the starting and ending points of the conductor (unit: km), measured using a Leica laser rangefinder and used to calculate line resistance. The conductor cross-sectional area parameter represents the geometric dimensions of the conductor cross section (unit: mm²).
[0099] In an embodiment of the present invention, line inventory data is called up through the power grid asset management system to extract conductor material attribute parameters; the actual ambient temperature value (unit: ° C) is collected in real time through a temperature and humidity sensor; the spatial length dimension (unit: km) is measured using a laser rangefinder; and the conductor cross-sectional area parameter (unit: mm²) is measured using a caliper.
[0100] Step 212: Calculating a reference resistivity value and a temperature drift compensation coefficient based on the conductor material property parameters and the measured ambient temperature value;
[0101] In this step, the reference resistivity value refers to the resistance per unit length of conductor at a standard temperature (20°C) (unit: Ω·mm² / m), obtained by consulting a material resistivity comparison table. The temperature drift compensation factor is the resistivity correction factor due to temperature changes, reflecting the effect of temperature rise on resistance.
[0102] In this embodiment of the present invention, a material resistivity database is queried and the material code is matched to obtain a reference resistivity value at 20°C. The temperature drift compensation coefficient α is calculated according to the formula: α = 1 + β × (T - 20), where β is the material temperature coefficient and T is the measured temperature (unit: °C).
[0103] Step 213: Calculating line geometric characteristic values according to the spatial length dimension and conductor cross-sectional area parameters;
[0104] In this step, the line geometric characteristic value refers to the ratio of length to cross-sectional area (unit: m -1 ), the physical meaning is the equivalent length per unit cross-sectional area.
[0105] In the embodiment of the present invention, the spatial length dimension is divided by the conductor cross-sectional area parameter to obtain a geometric characteristic value (unit: km / mm²), which is dimensionless but is used as a numerical value in subsequent calculations.
[0106] Step 214: multiplying the line geometric characteristic value and the temperature drift compensation coefficient to obtain a physical correction factor;
[0107] In this step, the physical correction factor refers to the correction coefficient that integrates geometric characteristics and temperature effects, and is used to compensate for changes in the physical state of the line.
[0108] In the embodiment of the present invention, the geometric characteristic value is multiplied by the temperature drift compensation coefficient to obtain a physical correction factor. For example, when the geometric characteristic value is 11666.67 and the temperature drift compensation coefficient is 1.062, the physical correction factor = 11666.67 × 1.062 = 12390.0, which is dimensionless.
[0109] Step 215: Calculate an electrical characteristic factor based on the dominant impedance component and the short-circuit capacity data of the grid connection point, and calculate a basic voltage sensitivity factor of each grid node in the access area of the virtual power plant in combination with the physical correction factor;
[0110] In this step, the electrical characteristic factor refers to the ratio of the dominant impedance component to the square root of the short-circuit capacity data, which is a dimensionless parameter that characterizes the strength of the power grid.
[0111] In the embodiment of the present invention, the electrical characteristic factor=dominant impedance component÷(square root of short-circuit capacity data); the basic sensitivity factor=electrical characteristic factor×physical correction factor.
[0112] The embodiments of the present invention solve the sensitivity drift caused by the temperature difference between day and night through temperature compensation; geometric correction eliminates the sensitivity attenuation of long-distance lines; and the fusion calculation of physical correction factors and electrical characteristic factors helps to improve the success rate of voltage regulation of subsequent reactive resource dynamic allocation strategies.
[0113] The present invention provides a specific embodiment, step 204, calculating a weight factor based on the real-time current carrying capacity and rated current carrying capacity parameters of the line associated with the dominant impedance component, specifically includes the following steps:
[0114] Step 221: Acquire the real-time current carrying capacity, rated current carrying parameters, and historical maximum current carrying records of the line associated with the dominant impedance component;
[0115] In this step, the historical maximum current record refers to the peak current (unit: A) of the line within the statistical period (usually 30 days). The maximum value is extracted based on the current data stored every 5 minutes by the data acquisition and monitoring system to reflect the extreme load capacity of the line.
[0116] In an embodiment of the present invention, the current carrying capacity (unit: A) is collected in real time through the current transformer on the line, the rated current carrying parameter (unit: A) is obtained by reading the power grid equipment ledger, and the maximum current carrying record (unit: A) in the past 30 days is retrieved from the historical database of the data acquisition and monitoring control system.
[0117] Step 222: Calculate the instantaneous approximation ratio of the real-time current carrying capacity to the rated current carrying capacity parameter, and determine the current carrying capacity fluctuation coefficient based on the historical maximum current carrying capacity record;
[0118] In this step, the instantaneous approach ratio is the ratio of the real-time current carrying capacity to the rated current carrying capacity parameter. This dimensionless quantity (ranging from 0 to 1.2) is used to quantify the current load factor. The current fluctuation coefficient is the ratio of the historical maximum current carrying capacity to the rated current carrying capacity multiplied by the fluctuation weight, reflecting the probability of line overload risk.
[0119] In the embodiment of the present invention, the instantaneous approximation ratio = real-time current carrying capacity ÷ rated current carrying parameter; the current carrying fluctuation coefficient = historical maximum current carrying record ÷ rated current carrying parameter × fluctuation weight, where the fluctuation weight is an empirical coefficient and takes a value of 0.7.
[0120] Step 223: Obtaining the conductor surface contamination level and environmental temperature and humidity parameters of the line associated with the dominant impedance component;
[0121] In this step, the conductor surface contamination level is determined based on the equivalent salt and ash density measurements, ranging from 0 to 4, corresponding to varying degrees of contamination severity. Ambient temperature and humidity parameters refer to the real-time temperature (°C) and relative humidity (%) at the line location, collected via a wireless sensor network and used to calculate heat dissipation conditions.
[0122] In this embodiment of the present invention, a UV imager carried by a drone is used to identify the contamination level of the conductor surface. For example, a salt density of 0.2 mg / cm² corresponds to a conductor surface contamination level of 3. A temperature and humidity sensor is used to collect the ambient temperature (unit: °C) and relative humidity (unit: %) to obtain the ambient temperature and humidity parameters.
[0123] Step 224: Determine the contamination thermal resistance increment coefficient based on the conductor surface contamination level, and calculate the air heat dissipation efficiency based on the ambient temperature and humidity parameters;
[0124] In this step, the contamination thermal resistance increment factor refers to the additional thermal resistance multiplier caused by contamination, reflecting the impact of contamination on temperature rise. Air cooling efficiency is a quantitative value of air cooling efficiency, and its physical meaning is the heat dissipation capacity per unit temperature rise.
[0125] In this embodiment of the present invention, a conductor surface contamination level mapping table is queried to determine the contamination thermal resistance increment coefficient. For example, a conductor surface contamination level 3 corresponds to a contamination thermal resistance increment coefficient of 1.25. Air cooling efficiency = inverse temperature / humidity compensation term, where humidity compensation term = 1 - 0.005 × H, where H is the relative humidity percentage. Air cooling efficiency = 1 / (T + A) × humidity compensation term, where A is a constant added to achieve temperature scale conversion.
[0126] Step 225: Calculate a dominant impedance reference value based on the dominant impedance component, and obtain a high-frequency harmonic distortion rate of the line associated with the dominant impedance component to calculate a skin effect depth correction coefficient;
[0127] In this step, the dominant impedance reference value refers to the dimensionless value of the dominant impedance component normalized by the system reference impedance, used to unify the dimensions for calculations. The high-frequency harmonic distortion rate (HFHD) refers to the percentage of the effective value of harmonic currents of the 13th order and above to the fundamental current, obtained through fast Fourier transform analysis. The skin effect depth correction factor (SDE) compensates for the reduction in effective cross-sectional area of the conductor caused by high-frequency currents and is used to correct AC resistance.
[0128] In an embodiment of the present invention, a power quality analyzer performs fast Fourier transform analysis on the current signal, and the high-frequency harmonic distortion rate is calculated based on the ratio of the square root of the sum of the squares of the effective values of the 13th to 50th harmonics to the fundamental current. The skin effect correction coefficient = 1 + 0.02 × high-frequency harmonic distortion rate.
[0129] Step 226: Multiplying the instantaneous approach ratio, the current fluctuation coefficient, the contamination thermal resistance increment coefficient, the air cooling efficiency, and the skin effect depth correction coefficient to obtain a comprehensive current carrying capacity factor, and combining the dominant impedance reference value to calculate a weighting factor;
[0130] In this step, the comprehensive current carrying capacity factor refers to a dimensionless composite coefficient that integrates the effects of load, pollution, heat dissipation, and harmonics and represents the line safety margin.
[0131] In the embodiment of the present invention, the comprehensive current carrying capacity factor = instantaneous approach ratio × current carrying fluctuation coefficient × pollution thermal resistance increment coefficient × air heat dissipation efficiency × skin effect depth correction coefficient, and the weighting factor = comprehensive current carrying capacity factor ÷ dominant impedance reference value.
[0132] The embodiment of the present invention dynamically associates the physical state of the line (such as pollution level, temperature and humidity) with electrical parameters (such as current carrying capacity fluctuations) to provide multi-dimensional dynamic input for the weight factor, so that the dynamic allocation strategy of reactive resources can respond to changes in grid load and environmental disturbances in real time, thereby improving the adaptability of the virtual power plant to the volatility of distributed power sources.
[0133] The present invention provides a specific embodiment, step 103, locating a weak line section associated with a grid node based on the real-time impedance spectrum measurement data, specifically comprising the following steps:
[0134] Step 301: Analyze the real-time impedance spectrum measurement data to obtain impedance amplitude and phase angle change;
[0135] In this step, the impedance amplitude refers to the modulus of the line impedance at the fundamental frequency (unit: Ω), reflecting the electrical strength of the line. The phase angle change refers to the difference between the current phase angle and the reference phase angle (unit: degrees). The reference value is the initial measurement of the line and is used to diagnose loose connections. A value greater than 5° indicates risk.
[0136] In an embodiment of the present invention, a digital signal processor is used to execute an operational process including frequency domain data extraction, complex impedance decomposition, impedance amplitude calculation, and phase angle change calculation. Specifically, a 1024-point fast Fourier transform is first performed on the real-time impedance spectrum measurement data in the 0.1-10kHz frequency band to locate the 50Hz fundamental frequency point. The complex impedance at the fundamental frequency point is extracted: the real part = the real part of the spectrum; the imaginary part = the imaginary part of the spectrum; the impedance amplitude = the square root (the square of the real part + the square of the imaginary part); the current phase angle = the inverse tangent function (imaginary part / real part) × (180 / π); and the phase angle change = the current phase angle - the historical reference phase angle.
[0137] Step 302: Calculate the fluctuation coefficient of the phase angle variation within a preset time window, and simultaneously calculate the deviation of the impedance amplitude from the historical reference amplitude, and mark the corresponding line segment where the deviation exceeds the impedance anomaly threshold and the fluctuation coefficient exceeds the preset stability threshold as a candidate weak segment;
[0138] In this step, the fluctuation coefficient refers to the absolute ratio of the standard deviation to the mean of the phase angle variation within a preset time window. A value > 2.0 indicates mechanical instability. The historical baseline amplitude refers to the 30-day moving average of the impedance amplitude when the line is in a healthy state (unit: Ω), calculated based on the historical database of the data acquisition and monitoring control system. The impedance anomaly threshold refers to the preset impedance deviation rate threshold (usually 20%), calculated as: |Current Amplitude - Historical Baseline| / Historical Baseline × 100%. The preset stability threshold refers to the safe limit of the phase angle fluctuation coefficient (usually 2.0). Exceeding this value is considered structural instability. Candidate weak sections are line sections that meet the dual threshold conditions. The data structure is [line ID, start point, end point] and the length is 0.5-2 km.
[0139] In an embodiment of the present invention, the fluctuation coefficient is calculated within a 15-minute time window, where the fluctuation coefficient = |standard deviation of the phase angle change ÷ mean of the phase angle change|; the deviation = (current impedance amplitude - historical reference amplitude) ÷ historical reference amplitude; when the deviation rate is greater than the impedance anomaly threshold and the fluctuation coefficient is greater than the preset stability threshold, the corresponding line segment is marked as a candidate weak segment.
[0140] Step 303: encoding the candidate weak section into a directed line segment according to the starting grid node and the ending grid node of the candidate weak section;
[0141] In this step, the starting and ending grid nodes refer to the topological endpoints of the line segment. The node numbering rule is voltage level-region-sequence number. A directed line segment is a directional representation of a line segment, defined as an ordered node pair [starting point, ending point].
[0142] In an embodiment of the present invention, based on the distribution network topology database, the starting grid node (such as NodeA) and the ending grid node (such as NodeB) of the candidate weak section are formed into an ordered pair to generate a directed segment. The specific data structure is [segment ID: L15, starting point: NodeA, end point: NodeB].
[0143] Step 304: Traverse all directed line segments. If the end node of the first directed line segment is the same as the start node of the second directed line segment, merge the first directed line segment and the second directed line segment into a target line segment. Combine all target line segments to generate a weak line segment.
[0144] In this step, the first and second directed line segments refer to adjacent line segment objects in the merge operation. The endpoint of the first directed line segment must be equal to the starting point of the second directed line segment. The target line segment refers to the continuous line segment formed after the merge, containing ≥ 2 nodes.
[0145] In this embodiment of the present invention, all first directed line segments are traversed. If the first directed line segment (such as the end point of L15 = NodeB) and the second directed line segment (such as the starting point of L16 = NodeB) have the same node, they are merged into a continuous target line segment [NodeA, NodeB, NodeC]. The recursive execution is performed until there are no node-overlapping line segments, and the final weak line segment is output.
[0146] The embodiments of the present invention solve the pain point that traditional methods are difficult to quickly identify key risk points in the power grid through real-time impedance spectrum data analysis and positioning of weak line sections, and achieve accurate positioning of vulnerable areas of the power grid; provide a clear optimization direction for the subsequent dynamic allocation of reactive resources, and enhance the emergency response capabilities of virtual power plants under extreme working conditions.
[0147] The present invention provides a specific embodiment, step 104, generating a reactive resource dynamic allocation strategy using a deep reinforcement learning agent based on the voltage sensitivity parameter and the information of the weak line section, specifically comprising the following steps:
[0148] Step 401: creating a feature group for each grid node, the feature group including node optimization sensitivity parameters and weak nodes, the weak nodes being determined based on whether the corresponding node is a starting grid node or an ending grid node of a weak line section;
[0149] In this step, a feature group is a vector container describing the electrical state of a grid node. Its data structure is {node ID, node optimization sensitivity parameter, weak node identifier}. The node optimization sensitivity parameter is modified by the risk enhancement factor. The node optimization sensitivity parameter is a dimensionless sensitivity value enhanced by the voltage offset integral and topology weight. The calculation formula is: optimized sensitivity = base sensitivity × (1 + voltage offset integral × connection weight), with a value range of 0.05-0.15. Weak nodes are grid nodes corresponding to the endpoints of weak line segments.
[0150] In an embodiment of the present invention, the connection weight coefficient of the starting grid node and the ending grid node is calculated based on the node connection relationship in the weak line section; the connection weight coefficient corresponding to the grid node marked as the weak node is multiplied by the voltage offset integral value to generate a risk enhancement factor, and the node optimization sensitivity parameter is calculated in combination with the voltage sensitivity parameter, and the node optimization sensitivity parameter is combined with the weak node to form a feature group core element, and the feature group core element is combined with the hierarchical position code of the grid node in the virtual power plant to generate a feature group for each grid node.
[0151] Step 402: Determine the target grid node to which each reactive resource device is connected according to a preset device node association table;
[0152] In this step, the preset device node association table refers to a predefined device topology database that stores the mapping between reactive resource device IDs and their physical connection nodes. The target grid node refers to the distribution network node ID to which the reactive resource device is directly connected, which is obtained by querying the device node association table.
[0153] In the embodiment of the present invention, a preset device node association table is read, a mapping relationship between reactive resource device IDs and grid nodes is analyzed, and a target grid node to which each reactive resource device is connected is determined.
[0154] Step 403: Analyze the feature group and the target grid node using a pre-trained deep reinforcement learning agent to generate a regulation command value for each reactive resource device;
[0155] In this step, the adjustment command value refers to the normalized control command value (in the range [-1,1]) output by the deep reinforcement learning agent. A positive number indicates the emission of reactive power / discharge, and a negative number indicates the absorption of reactive power / charging.
[0156] In an embodiment of the present invention, the feature group is sorted into a 128-dimensional vector by node number, and the target node information is encoded as a 32-dimensional one-hot vector; a 160-dimensional state vector is input through a pre-trained deep reinforcement learning agent, and an adjustment instruction value of the continuous action space [-1,1] is output, wherein the action output layer uses a hyperbolic tangent activation function to ensure that the instruction value is in the [-1,1] interval.
[0157] Step 404: converting the regulation command value into a reactive power setting value of the distributed power source and a charge / discharge power setting value of the energy storage converter, so as to generate a reactive resource dynamic allocation strategy according to the reactive power setting value and the charge / discharge power setting value;
[0158] In this step, the reactive power setpoint refers to the actual control parameter of the distributed generation (kVar), calculated by multiplying the adjustment command value by the rated capacity, and is in the range [-rated value, +rated value]. The charge / discharge power setpoint refers to the active power command of the energy storage converter (kW), calculated by multiplying the adjustment command value by the rated power. Negative values indicate charging.
[0159] In an embodiment of the present invention, the reactive power setting value of the distributed power supply = the adjustment instruction value × the rated capacity of the equipment; the charging and discharging power setting value of the energy storage converter = the adjustment instruction value × the rated power; and finally, the two setting values are encapsulated into a reactive resource dynamic allocation strategy.
[0160] The embodiment of the present invention generates a dynamic allocation strategy for reactive resources through a deep reinforcement learning intelligent agent, breaking through the limitations of traditional heuristic algorithms in real-time and adaptability, improving the real-time and coordination of multi-objective optimization scheduling of virtual power plants, and meeting the dynamic balance needs after a high proportion of renewable energy is connected.
[0161] The present invention provides a specific embodiment, step 401, creating a feature group for each power grid node, the feature group including a node optimization sensitivity parameter and a weak node, the weak node being determined based on whether the corresponding node is a starting power grid node or an ending power grid node of a weak line section, specifically comprising the following steps:
[0162] Step 411: collecting voltage fluctuation trajectory data of the grid node within a preset time window to calculate the voltage offset integral value;
[0163] In this step, voltage fluctuation trajectory data refers to a continuous sampling sequence (unit: pu) of grid node voltage within a preset time window (15 minutes). This data is collected at a rate of ≥30 frames per second using a phasor measurement device and reflects the transient voltage fluctuation characteristics. The voltage offset integral value refers to the cumulative deviation of the voltage from the reference value (unit: pu·s), which quantifies the severity of the voltage overshoot.
[0164] In this embodiment of the present invention, a phasor measurement device collects voltage fluctuation trajectory data (unit: pu) at power grid nodes within a preset 15-minute time window at a rate of 50 frames per second. The voltage offset integral value is calculated using the trapezoidal integration method: voltage offset integral value = ∑[|actual voltage - reference voltage| × time interval].
[0165] Step 412: Calculate the connection weight coefficients of the starting grid node and the ending grid node based on the node connection relationship in the weak line section;
[0166] In this step, node connectivity refers to the topological connectivity structure of grid nodes within a weak line segment, represented as an adjacency matrix generated using a graph theory algorithm. The connection weight coefficient is an indicator (0–1) of the importance of a node within the weak segment, reflecting the risk of fault propagation.
[0167] In an embodiment of the present invention, a node adjacency matrix is extracted from the weak line section topology, and a degree centrality algorithm is used to calculate a connection weight coefficient, where the connection weight coefficient = (the number of lines connected to the current node) / (the total number of lines in the weak section).
[0168] Step 413: Multiplying the connection weight coefficient corresponding to the grid node marked as a weak node by the voltage offset integral value to generate a risk enhancement factor, and combining the voltage sensitivity parameter to calculate the node optimization sensitivity parameter;
[0169] In this step, the risk enhancement factor refers to the risk gain coefficient (≥0) for weak nodes only, which is obtained by multiplying the connection weight coefficient and the voltage offset integral value, and is used to increase the sensitivity weight of high risks.
[0170] In the embodiment of the present invention, the risk enhancement factor is calculated only for weak nodes (i.e., the mark = 1); the risk enhancement factor = the connection weight coefficient × the voltage offset integral value; the node optimization sensitivity parameter = the voltage sensitivity parameter × (1 + the risk enhancement factor), and the non-weak nodes retain the original value.
[0171] Step 414: combining the node optimization sensitivity parameters and the weak nodes into a feature group core element, and combining the feature group core element with the hierarchical position code of the grid node in the virtual power plant to generate a feature group for each grid node;
[0172] In this step, the core element of the feature group refers to the core data unit that describes the electrical state of the node. The data structure is a two-tuple [node optimization sensitivity parameter, weak node]. The hierarchical position encoding refers to the hierarchical identification of the node in the virtual power plant control architecture and is used to distinguish control priorities.
[0173] In an embodiment of the present invention, the core elements of the feature group are encapsulated into a binary group [node optimization sensitivity parameter, weak node]; the grid nodes are assigned codes according to the three-level architecture of the virtual power plant, where the code of the main network access layer is 100, the code of the distribution network hub layer is 200, and the code of the user access layer is 300; and finally, the core elements of the feature group and the hierarchical codes are combined into a structured feature group [sensitivity value, weak node, hierarchical code].
[0174] The embodiments of the present invention solve the problem of insufficient representation of key characteristics of power grid nodes in traditional scheduling methods; the deep reinforcement learning agent can more comprehensively capture the topological characteristics and operating status of the power grid and generate more accurate reactive resource regulation instruction values; it improves the adaptive scheduling capability of virtual power plants in complex power grid environments and reduces the system stability risks caused by the volatility of distributed power sources.
[0175] Figure 2 The present invention provides a structural diagram of a multi-objective optimization scheduling system for a virtual power plant integrating deep reinforcement learning, such as Figure 2 As shown, the system includes:
[0176] An acquisition module 21 is used to acquire short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point;
[0177] A calculation module 22 is configured to calculate a voltage sensitivity parameter of each grid node in the access area of the virtual power plant based on the short-circuit capacity data;
[0178] a positioning module 23 for locating a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data;
[0179] The generation module 24 is used to generate a dynamic reactive resource allocation strategy based on the voltage sensitivity parameters and the information of the weak line section using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the equipment regulation loss.
[0180] The control module 25 is used to coordinately control the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter according to the reactive resource dynamic allocation strategy, so as to achieve multi-objective optimization scheduling of the virtual power plant.
[0181] Figure 2 The virtual power plant multi-objective optimization scheduling system integrating deep reinforcement learning can be executed Figure 1 The implementation principle and technical effects of the multi-objective optimization scheduling method for a virtual power plant integrated with deep reinforcement learning described in the illustrated embodiment will not be repeated here. The specific manner in which each module and unit performs operations in the multi-objective optimization scheduling system for a virtual power plant integrated with deep reinforcement learning in the above embodiment has been described in detail in the embodiments of the method and will not be elaborated on here.
[0182] In one possible design, Figure 2 The multi-objective optimization scheduling system of a virtual power plant integrating deep reinforcement learning in the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0183] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0184] The processing component 32 is used for the above Figure 1 The embodiment provides a multi-objective optimization scheduling method for a virtual power plant that integrates deep reinforcement learning.
[0185] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0186] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0187] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0188] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0189] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0190] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0191] The embodiment of the present invention further provides a computer storage medium storing a computer program, which can achieve the above-mentioned Figure 1 The illustrated embodiment provides a multi-objective optimization scheduling method for a virtual power plant that integrates deep reinforcement learning.
[0192] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0193] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0194] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multi-objective optimization scheduling method for virtual power plants integrating deep reinforcement learning, characterized in that: include: Obtain short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point; Calculating voltage sensitivity parameters of each grid node in the access area of the virtual power plant based on the short-circuit capacity data; locating a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data; Based on the voltage sensitivity parameter and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize voltage regulation compensation benefits and minimize equipment regulation losses; According to the reactive resource dynamic allocation strategy, the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter are coordinated and controlled to achieve multi-objective optimization scheduling of the virtual power plant.
2. The method according to claim 1, characterized in that Calculating voltage sensitivity parameters of each grid node in the access area of the virtual power plant based on the short-circuit capacity data includes: Calculate the cumulative value of the line impedance from the virtual power plant's grid connection point to each grid node in the virtual power plant's access area; Extracting a target impedance component whose proportion exceeds a set threshold value from the accumulated value of the line impedance, and using the target impedance component as a dominant impedance component; Calculating a basic voltage sensitivity factor of each grid node in the access area of the virtual power plant based on the dominant impedance component and the short-circuit capacity data of the grid connection point; Calculating a weighting factor based on the real-time current carrying capacity and the rated current carrying capacity parameters of the line associated with the dominant impedance component; The basic voltage sensitivity factor is superimposed on the weight factor to generate the voltage sensitivity parameter of each grid node in the access area of the virtual power plant.
3. The method according to claim 2, characterized in that Calculating the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant according to the dominant impedance component and the short-circuit capacity data of the grid connection point, including: Obtaining conductor material property parameters, ambient temperature measured value, spatial length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component; Calculating a reference resistivity value and a temperature drift compensation coefficient based on the conductor material property parameters and the measured ambient temperature value; Calculating line geometric characteristic values according to the spatial length dimension and conductor cross-sectional area parameters; Performing a product operation on the line geometric characteristic value and the temperature drift compensation coefficient to obtain a physical correction factor; The electrical characteristic factor is calculated based on the dominant impedance component and the short-circuit capacity data of the grid connection point, and combined with the physical correction factor, the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant is calculated.
4. The method according to claim 2, characterized in that A weighting factor is calculated based on the real-time current carrying capacity and rated current carrying capacity parameters of the line associated with the dominant impedance component, including: Obtaining the real-time current carrying capacity, rated current carrying capacity parameters and historical maximum current carrying capacity records of the line associated with the dominant impedance component; Calculating the instantaneous approximation ratio of the real-time current carrying capacity to the rated current carrying capacity parameter, and determining the current carrying capacity fluctuation coefficient based on the historical maximum current carrying capacity record; Obtaining the conductor surface contamination level and environmental temperature and humidity parameters of the line associated with the dominant impedance component; According to the contamination level of the conductor surface, determine the contamination thermal resistance increment coefficient, and calculate the air heat dissipation efficiency in combination with the ambient temperature and humidity parameters; Calculating a dominant impedance reference value based on the dominant impedance component, and obtaining a high-frequency harmonic distortion rate of the line associated with the dominant impedance component to calculate a skin effect depth correction coefficient; The instantaneous approach ratio, the current fluctuation coefficient, the pollution thermal resistance increment coefficient, the air heat dissipation efficiency and the skin effect depth correction coefficient are multiplied to obtain a comprehensive current carrying capacity factor, and the weight factor is calculated in combination with the dominant impedance reference value.
5. The method according to claim 1, characterized in that Locating a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data, comprising: Analyzing the real-time impedance spectrum measurement data to obtain impedance amplitude and phase angle change; Calculating the fluctuation coefficient of the phase angle variation within a preset time window, and simultaneously calculating the deviation of the impedance amplitude from a historical reference amplitude, and marking the corresponding line segment where the deviation exceeds an impedance anomaly threshold and the fluctuation coefficient exceeds a preset stability threshold as a candidate weak segment; encoding the candidate weak section into a directed line segment according to the starting grid node and the ending grid node of the candidate weak section; Traverse all directed line segments. If the end node of the first directed line segment is the same as the start node of the second directed line segment, merge the first directed line segment and the second directed line segment into a target line segment. Combine all target line segments to generate a weak line segment.
6. The method according to claim 1, characterized in that Based on the voltage sensitivity parameter and the information of the weak line section, a dynamic reactive resource allocation strategy is generated using a deep reinforcement learning agent, including: Creating a feature group for each grid node, the feature group comprising a node optimization sensitivity parameter and a weak node, the weak node being determined based on whether the corresponding node is a starting grid node or an ending grid node of a weak line segment; Determine the target grid node to which each reactive resource device is connected according to a preset device node association table; Utilizing a pre-trained deep reinforcement learning agent, the feature group and the target grid node are analyzed to generate a regulation instruction value for each reactive resource device; The regulation command value is converted into a reactive power setting value of the distributed power source and a charge and discharge power setting value of the energy storage converter, so as to generate a reactive resource dynamic allocation strategy according to the reactive power setting value and the charge and discharge power setting value.
7. The method according to claim 6, characterized in that Creating a feature group for each grid node, the feature group including node optimization sensitivity parameters and weak nodes, wherein the weak nodes are determined based on whether the corresponding node is a starting grid node or an ending grid node of a weak line section, including: Collect voltage fluctuation trajectory data of grid nodes within a preset time window to calculate the voltage offset integral value; Calculate the connection weight coefficients of the starting grid node and the ending grid node based on the node connection relationship in the weak line section; The connection weight coefficient corresponding to the grid node marked as a weak node is multiplied by the voltage offset integral value to generate a risk enhancement factor, and the node optimization sensitivity parameter is calculated in combination with the voltage sensitivity parameter; The node optimization sensitivity parameters and the weak nodes are combined into a feature group core element, and the feature group core element is combined with the hierarchical position code of the power grid node in the virtual power plant to generate a feature group for each power grid node.
8. A virtual power plant multi-objective optimization scheduling system integrating deep reinforcement learning, characterized by: include: An acquisition module is used to obtain short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant's grid connection point; a calculation module, configured to calculate a voltage sensitivity parameter of each grid node in an access area of the virtual power plant based on the short-circuit capacity data; a positioning module, configured to locate a weak line section associated with a power grid node based on the real-time impedance spectrum measurement data; a generation module for generating a dynamic reactive resource allocation strategy based on the voltage sensitivity parameter and information about the weak line section using a deep reinforcement learning agent, wherein the deep reinforcement learning agent aims to maximize voltage regulation compensation benefits and minimize equipment regulation losses; The control module is used to coordinately control the reactive output power of the distributed power source and the charging and discharging power of the energy storage converter according to the dynamic allocation strategy of reactive resources, so as to achieve multi-objective optimization scheduling of the virtual power plant.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a virtual power plant multi-objective optimization scheduling method integrating deep reinforcement learning as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual power plant dynamic regulation and control method based on deep reinforcement learning
CN119543316A
Multi-target scheduling optimization method and system for virtual power plant to participate in electricity market
CN120087801A