Virtual power plant multi-objective optimization scheduling method and system fusing deep reinforcement learning
By obtaining the connection point data and real-time impedance spectrum of the virtual power plant, combined with deep reinforcement learning agents, identifying weak lines and generating dynamic allocation strategies for reactive resource, the multi-objective optimization scheduling problem of virtual power plants in complex power grid environments is solved, and the voltage stability and economy of the power grid are improved.
Patent Information
- Application Number
- CN202510858341.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the face of complex and changing power grid environment, existing virtual power plant scheduling methods are difficult to achieve multi-objective optimization scheduling, especially in terms of voltage stability control and dynamic allocation of reactive resources, there are problems of model mismatch, response lag and distributed resource dispersion.
By obtaining the short-circuit capacity data and real-time impedance spectrum measurement data of the network connection points of the virtual power plant, the voltage sensitivity parameters of the power grid node are calculated, and the deep reinforcement learning agent is used to locate weak line segments to generate a dynamic allocation strategy for reactive resource, and the goal of maximizing the profit of voltage regulation compensation and minimizing equipment regulation loss, the power of distributed power and energy storage converters is coordinated.
It realizes rapid collaborative optimization scheduling of virtual power plants in complex power grid environments, improves grid voltage stability and operational economy, meets the requirements of real-time and coordination, and overcomes the problems of low search efficiency and adaptive adjustment of traditional methods in high-dimensional and nonlinear environments.
Smart Images

Figure CN120377361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual power plant dispatching, and particularly to a multi-objective optimal dispatching method and system for virtual power plants integrating deep reinforcement learning. Background Art
[0002] With the continuous transformation of the energy structure and the continuous increase in the penetration rate of renewable energy, virtual power plants, as an important technical means to integrate distributed energy resources and improve the flexibility of power grid operation, are gradually playing a key role in the power system. In practical applications, virtual power plants need to achieve multi-objective optimal dispatching in a complex and changing power grid environment, especially facing severe challenges in voltage stability control and reactive power resource dynamic allocation. Due to the scattered access locations and strong volatility of distributed power sources, traditional centralized dispatching methods are difficult to meet the requirements of real-time performance and coordination.
[0003] A currently typical solution is a method combining model predictive control and heuristic algorithms. This solution constructs a mathematical model of each unit inside the virtual power plant and uses heuristic algorithms such as genetic algorithms or particle swarm optimization to solve the optimal reactive power allocation scheme during the rolling optimization process. However, the existing solutions still have obvious limitations. For example, they rely on accurate system modeling, and after the access of distributed energy, the power grid topology changes frequently, resulting in prominent model mismatch problems and affecting control accuracy; the search efficiency of heuristic algorithms in high-dimensional state spaces is low, and it is difficult to achieve real-time response and online adaptive adjustment, etc. Summary of the Invention
[0004] The present invention provides a multi-objective optimal dispatching method and system for virtual power plants integrating deep reinforcement learning to solve the problems of prominent model mismatch and insufficient dispatching accuracy in the prior art; it is difficult to achieve real-time response and online adaptive adjustment, etc.
[0005] In a first aspect, the present invention provides a multi-objective optimal dispatching method for virtual power plants integrating deep reinforcement learning, including: Obtaining the short-circuit capacity data and real-time impedance spectrum measurement data of the connection point of the virtual power plant; Calculating the voltage sensitivity parameters of each power grid node in the access area of the virtual power plant based on the short-circuit capacity data; Locating the weak line sections associated with the power grid nodes according to the real-time impedance spectrum measurement data; Generating a reactive power resource dynamic allocation strategy by using a deep reinforcement learning agent according to the voltage sensitivity parameters and the information of the weak line sections, where the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss; According to the reactive power resource dynamic allocation strategy, coordinately control the reactive power output power of distributed power sources and the charging and discharging power of energy storage converters to achieve multi-objective optimal scheduling of the virtual power plant.
[0006] Optionally, based on the short-circuit capacity data, calculate the voltage sensitivity parameters of each grid node in the access area of the virtual power plant, including: Calculate the cumulative value of the line impedance from the grid connection point of the virtual power plant to each grid node in the access area of the virtual power plant; Extract the target impedance component in the cumulative value of the line impedance whose proportion exceeds the set threshold, and use the target impedance component as the dominant impedance component; According to the dominant impedance component and the short-circuit capacity data of the grid connection point, calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant; Calculate the weight factor according to the real-time current-carrying capacity and rated current-carrying parameters of the line associated with the dominant impedance component; Superimpose the basic voltage sensitivity factor and the weight factor to generate the voltage sensitivity parameters of each grid node in the access area of the virtual power plant.
[0007] Optionally, according to the dominant impedance component and the short-circuit capacity data of the grid connection point, calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant, including: Obtain the conductor material property parameters, measured ambient temperature value, spatial length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component; Calculate the reference resistivity value and temperature drift compensation coefficient according to the conductor material property parameters and the measured ambient temperature value; Calculate the line geometric characteristic value according to the spatial length dimension and the conductor cross-sectional area parameters; Perform a multiplication operation on the line geometric characteristic value and the temperature drift compensation coefficient to obtain the physical correction factor; Calculate the electrical characteristic factor according to the dominant impedance component and the short-circuit capacity data of the grid connection point, and combine the physical correction factor to calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant.
[0008] Optionally, calculate the weight factor according to the real-time current-carrying capacity and rated current-carrying parameters of the line associated with the dominant impedance component, including: Obtain the real-time current-carrying capacity, rated current-carrying parameters, and historical maximum current-carrying record of the line associated with the dominant impedance component; Calculate the instantaneous approximation ratio of the real-time current-carrying capacity to the rated current-carrying parameters, and at the same time, based on the historical maximum current-carrying record, determine the current-carrying fluctuation coefficient; Obtain the conductor surface pollution level and environmental temperature and humidity parameters of the line associated with the dominant impedance component; Determine the pollution thermal resistance increment coefficient according to the conductor surface pollution level, and calculate the air heat dissipation efficiency in combination with the environmental temperature and humidity parameters; Calculate the dominant impedance reference value according to the dominant impedance component, and obtain the high-frequency harmonic distortion rate of the line associated with the dominant impedance component to calculate the skin effect depth correction coefficient; Perform a product operation on the instantaneous approximation ratio, the current-carrying fluctuation coefficient, the pollution thermal resistance increment coefficient, the air heat dissipation efficiency, and the skin effect depth correction coefficient to obtain a comprehensive current-carrying capacity factor, and calculate a weight factor in combination with the dominant impedance reference value.
[0009] Optionally, according to the real-time impedance spectrum measurement data, locate the weak line section associated with the power grid node, including: Analyze the real-time impedance spectrum measurement data to obtain the impedance amplitude and phase angle change amount; Calculate the fluctuation coefficient of the phase angle change amount within a preset time window, and at the same time calculate the deviation amount between the impedance amplitude and the historical reference amplitude. Mark the corresponding line section where the deviation amount exceeds the impedance anomaly threshold and the fluctuation coefficient exceeds the preset stability threshold as a candidate weak section; Encode the candidate weak section as a directed line segment according to the starting power grid node and the ending power grid node of the candidate weak section; Traverse all directed line segments. If the ending node of the first directed line segment is the same as the starting node of the second directed line segment, then merge the first directed line segment and the second directed line segment into a target line segment, and generate a weak line section in combination with all target line segments.
[0010] Optionally, according to the voltage sensitivity parameter and the information of the weak line section, use a deep reinforcement learning agent to generate a reactive power resource dynamic allocation strategy, including: Create a feature group for each power grid node. The feature group includes a node optimization sensitivity parameter and a weak node, and the weak node is determined according to whether the corresponding node is the starting power grid node or the ending power grid node of the weak line section; Determine the target power grid node connected to each reactive power resource device according to the preset device node association table; Use the pre-trained deep reinforcement learning agent to analyze the feature group and the target power grid node to generate the adjustment command value of each reactive power resource device; Convert the adjustment command value into a reactive power setting value of the distributed power source and a charge and discharge power setting value of the energy storage converter to generate a reactive power resource dynamic allocation strategy according to the reactive power setting value and the charge and discharge power setting value.
[0011] Optionally, a feature group is created for each grid node, and the feature group includes a node optimization sensitivity parameter and a weak node, and the weak node is determined according to whether the corresponding node is the starting grid node or the ending grid node of a weak line section, including: Collect the voltage fluctuation trajectory data of the grid node within a preset time window to calculate the voltage offset integral value; According to the node connection relationship in the weak line section, calculate the connection weight coefficient between the starting grid node and the ending grid node; Multiply the connection weight coefficient corresponding to the grid node marked as a weak node by the voltage offset integral value to generate a risk enhancement factor, and combine it with the voltage sensitivity parameter to calculate the node optimization sensitivity parameter; Combine the node optimization sensitivity parameter and the weak node into the core elements of the feature group, and combine the core elements of the feature group with the hierarchical position coding of the grid node in the virtual power plant to generate the feature group of each grid node.
[0012] In a second aspect, the present invention provides a virtual power plant multi-objective optimization scheduling system integrating deep reinforcement learning, including: An acquisition module for acquiring the short-circuit capacity data and real-time impedance spectrum measurement data of the connection point of the virtual power plant; A calculation module for calculating the voltage sensitivity parameters of each grid node in the access area of the virtual power plant based on the short-circuit capacity data; A positioning module for positioning the weak line section associated with the grid node according to the real-time impedance spectrum measurement data; A generation module for generating a reactive power resource dynamic allocation strategy by using a deep reinforcement learning agent according to the voltage sensitivity parameter and the information of the weak line section, and the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss.
[0013] A control module for coordinately controlling the reactive power output power of the distributed power source and the charge and discharge power of the energy storage converter according to the reactive power resource dynamic allocation strategy to achieve the multi-objective optimization scheduling of the virtual power plant.
[0014] In a third aspect, an embodiment of the present invention provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a virtual power plant multi-objective optimization scheduling method as described in the first aspect above.
[0015] Fourthly, an embodiment of the present invention provides a computer storage medium storing a computer program, which, when executed by a computer, implements a multi-objective optimal scheduling method for a virtual power plant integrating deep reinforcement learning as described in the first aspect.
[0016] In the present invention, short-circuit capacity data and real-time impedance spectrum measurement data of the connection point of the virtual power plant are obtained; based on the short-circuit capacity data, voltage sensitivity parameters of each power grid node in the access area of the virtual power plant are calculated; according to the real-time impedance spectrum measurement data, weak line sections associated with the power grid nodes are located; according to the voltage sensitivity parameters and the information of the weak line sections, a reactive power resource dynamic allocation strategy is generated by using a deep reinforcement learning agent, and the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss; according to the reactive power resource dynamic allocation strategy, the reactive power output power of the distributed power source and the charge and discharge power of the energy storage converter are coordinately controlled to achieve multi-objective optimal scheduling of the virtual power plant. The technical solution provided by the present invention breaks through the dependence of traditional methods on a fixed power grid model, perceives the dynamic topology change and impedance characteristics of the power grid in real time, provides a dynamic data basis for subsequent calculations to adapt to the fluctuations of renewable energy, and solves the problem of model mismatch caused by the access of distributed power sources; quantifies the sensitivity of the voltage of each node in the power grid to the change of reactive power, accurately identifies the key nodes with high voltage regulation efficiency, and provides a scientific basis for the optimal allocation of resources, overcoming the defect that traditional centralized scheduling is difficult to handle the dispersion and volatility of the positions of distributed power sources; actively identifies the vulnerable links (i.e., weak line sections) in the power grid that are prone to voltage instability or power quality problems, focuses the optimization attention on the risk areas, improves the overall operation robustness of the system, and makes up for the shortcoming of the traditional method's insufficient perception of local power grid risks; without the premise of an accurate mathematical model, the agent learns online by interacting with the environment, efficiently coordinates the dual objectives of maximizing the voltage regulation compensation benefit and minimizing the device regulation loss, and solves the core problems of low search efficiency, difficulty in real-time response and adaptive adjustment of heuristic algorithms in a high-dimensional, non-linear and strongly uncertain environment; converts the optimization strategy into direct and coordinated control instructions for the reactive power of the distributed power source and the energy storage converter, realizes the fast aggregation response and collaborative optimal scheduling of the dispersed resources in the virtual power plant, improves the voltage stability and operation economy of the power grid, and meets the real-time and collaborative requirements in a complex and changeable power grid environment.Further, calculate the voltage offset integral value based on the voltage fluctuation trajectory of the grid nodes, and calculate the risk enhancement factor by combining the connection weights of the nodes in the weak line section; secondly, fuse this factor with the voltage sensitivity parameter to generate a node optimization sensitivity parameter that can better reflect the optimization value and risk; then, combine the node optimization sensitivity parameter, the state of the weak nodes, and the hierarchical position coding of the nodes in the virtual power plant to form a feature group for each node; at the same time, according to the preset equipment-node association table, determine the target grid nodes connected to each reactive power resource device; finally, use the pre-trained deep reinforcement learning agent to analyze these feature groups and target grid nodes, directly output the optimal adjustment instruction value for each reactive power resource device, and then convert it into a specific power setting value to form a dynamic allocation strategy, enabling the reinforcement learning agent to deeply understand the grid dynamic characteristics and control priorities in an environment where the model parameters are unknown or changing, overcoming the problem of reduced control accuracy caused by model mismatch and topological changes in traditional methods; transform the high-dimensional and complex multi-objective optimization problem into an intelligent agent decision-making process based on rich features, avoid the inefficient search of heuristic algorithms in high-dimensional spaces, and greatly improve the real-time performance and computational efficiency of generating high-quality and executable strategies in a highly uncertain environment, meeting the urgent need for fast adaptive scheduling of virtual power plants.
[0017] These aspects or other aspects of the present invention will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 Shows a flowchart of a multi-objective optimization scheduling method for a virtual power plant integrating deep reinforcement learning provided by the present invention; Figure 2 Shows a schematic structural diagram of a multi-objective optimization scheduling system for a virtual power plant integrating deep reinforcement learning provided by the present invention; Figure 3 Shows a schematic structural diagram of a computing device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.
[0021] In some of the processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are different types.
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present invention.
[0023] Aiming at the problems of the existing virtual power plant optimal scheduling method based on model predictive control and heuristic algorithm, such as frequent changes in the power grid topology, decreased system modeling accuracy, and low search efficiency in the high-dimensional state space, it is difficult to meet the real-time and collaborative requirements of voltage stability control and reactive power resource dynamic allocation in a complex and changeable environment. The present invention proposes a virtual power plant multi-objective optimal scheduling method integrating deep reinforcement learning, making the scheduling no longer rely on accurate mathematical modeling, but by introducing a deep reinforcement learning agent, combining the short-circuit capacity data and real-time impedance spectrum measurement data of the virtual power plant connection point, extracting key voltage sensitivity parameters and identifying weak line sections, so as to construct an adaptive decision-making mechanism guided by maximizing the voltage regulation compensation benefit and minimizing the device regulation loss. By online learning the power grid operation state and environmental changes, the deep reinforcement learning agent can achieve the coordinated control of the reactive power output of distributed power sources and the charging and discharging power of energy storage converters, improve the response speed and multi-objective optimization ability of the scheduling strategy, overcome the limitations of traditional methods in dynamic adaptability and computational efficiency, and meet the urgent needs of the virtual power plant for intelligent and real-time scheduling in the new power system.
[0024] Figure 1 A flowchart of a virtual power plant multi-objective optimal scheduling method integrating deep reinforcement learning is provided for the embodiments of the present invention, as Figure 1 shown, the method includes: Step 101: Obtain the short-circuit capacity data and real-time impedance spectrum measurement data of the connection point of the virtual power plant; In this step, the virtual power plant refers to a cloud control system that aggregates distributed power sources, energy storage systems, and controllable loads, including photovoltaic inverters, energy storage converters, and load controllers, and realizes wide-area coordinated control through a communication network. The connection point refers to the physical connection point between the virtual power plant and the public power grid, usually the busbar of a 10kV / 35kV substation, and voltage transformers and current transformers are configured for electrical parameter measurement. The short-circuit capacity data refers to the apparent power value (unit: MVA) when the three-phase short circuit occurs at the connection point, which is calculated by a power system analysis comprehensive program or a power system analysis software package and reflects the reciprocal of the equivalent impedance of the power grid. The real-time impedance spectrum measurement data refers to the set of impedance amplitude (unit: Ω) and phase angle (unit: °) in the frequency band of 0.1 - 10kHz, which is obtained by the frequency scanning method and contains 512 frequency point data.
[0025] In the embodiment of the present invention, the short-circuit capacity data of the connection point of the virtual power plant is collected through the data acquisition and monitoring system of the power grid dispatching center, and this data comes from the power system short-circuit calculation database; at the same time, a distributed frequency response analyzer is used to perform frequency sweep excitation (0.1 - 10kHz) on the connection point, and the real-time impedance spectrum measurement data is obtained.
[0026] Step 102: Based on the short-circuit capacity data, calculate the voltage sensitivity parameters of each power grid node in the access area of the virtual power plant; In this step, the access area refers to the distribution network range managed by the virtual power plant, with a radial or ring network topology, including 35 power grid nodes and 28 feeders. A power grid node refers to a line connection point or equipment access point in the distribution network, with a unique number, and an intelligent meter is configured to monitor electrical quantities. The voltage sensitivity parameter refers to the ratio of the node voltage change rate to the reactive power injection change amount (unit: V / kVar), and the numerical range is 0.02 - 0.15.
[0027] In the embodiment of the present invention, first, the numerical quantity of the short-circuit capacity data is parsed, and combined with the target impedance component in the cumulative value of the line impedance from the connection point of the virtual power plant to each power grid node in the access area of the virtual power plant that exceeds the set threshold, the basic voltage sensitivity factor of each power grid node in the access area of the virtual power plant is calculated; according to the real-time current-carrying capacity and rated current-carrying parameters of the line associated with the dominant impedance component, the weight factor is calculated and superimposed with the basic voltage sensitivity factor to generate the voltage sensitivity parameters of each power grid node in the access area of the virtual power plant.
[0028] Step 103: Locate the weak line section associated with the power grid node according to the real-time impedance spectrum measurement data; In this step, the weak line section refers to a physical line section where the impedance abnormally increases > 20%, which is defined by the starting and ending power grid node coordinates, and the length is usually 0.5 - 3km.
[0029] In the embodiments of the present invention, fast Fourier transform is performed on the real-time impedance spectrum measurement data to extract the fundamental wave impedance amplitude and the phase angle change amount. The fluctuation coefficient of the phase angle change amount within a preset time window is calculated through a sliding time window, and the deviation amount between the impedance amplitude and the historical reference amplitude is compared. The corresponding line segment with the deviation amount exceeding the impedance anomaly threshold and the fluctuation coefficient exceeding the preset stability threshold is marked as a candidate weak section. Based on the graph theory algorithm, the target line segments in all candidate weak sections are merged according to the node connection relationship to generate a weak line section.
[0030] Step 104: According to the voltage sensitivity parameter and the information of the weak line section, use a deep reinforcement learning agent to generate a reactive power resource dynamic allocation strategy, where the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss; In this step, the reactive power resource dynamic allocation strategy refers to a set of control instructions output by the deep reinforcement learning agent, and the data structure is in JSON format.
[0031] In the embodiments of the present invention, a feature group is created for each power grid node. According to the preset device node association table, the target power grid node connected to each reactive power resource device is determined; a pre-trained deep reinforcement learning agent is used to analyze the feature group and the target power grid node to generate the adjustment instruction value of each reactive power resource device, and convert it into the reactive power setting value of the distributed power source and the charge and discharge power setting value of the energy storage converter. According to the reactive power setting value and the charge and discharge power setting value, a reactive power resource dynamic allocation strategy is generated.
[0032] Step 105: According to the reactive power resource dynamic allocation strategy, coordinately control the reactive power output power of the distributed power source and the charge and discharge power of the energy storage converter to achieve the multi-objective optimal scheduling of the virtual power plant; In this step, the reactive power output power refers to the capacitive or inductive reactive power provided by the distributed power source (unit: kVar), and the adjustment range of the photovoltaic inverter is -100% to +100% of the rated capacity. The charge and discharge power refers to the active power instruction of the energy storage converter (unit: kW), a positive value indicates discharging, a negative value indicates charging, and the limit is ±1C rate.
[0033] In the embodiments of the present invention, the setting value instruction in the reactive power resource dynamic allocation strategy is parsed, and the reactive power setting value (unit: kVar) is sent to the photovoltaic inverter through the substation communication network and system, and the charge and discharge power instruction (unit: kW) is sent to the energy storage converter through Modbus TCP. Among them, a master-slave control architecture is adopted to ensure the timing synchronization of the power instructions of multiple devices, and the voltage at the virtual power plant grid connection point is stabilized in the range of 0.95 - 1.05 pu.
[0034] The embodiments of the present invention solve the complex problems of inaccurate adjustment of resource allocation, frequent device actions, and insufficient economy in traditional methods, and achieve the comprehensive optimization goal of voltage qualification rate and device loss.
[0035] The present invention provides a specific embodiment. Step 102: Based on the short-circuit capacity data, calculate the voltage sensitivity parameters of each power grid node in the access area of the virtual power plant, which specifically includes the following steps: Step 201: Calculate the cumulative value of line impedances from the grid connection point of the virtual power plant to each power grid node in the access area of the virtual power plant; In this step, the cumulative value of line impedances refers to the arithmetic sum of the modulus values of all line impedances on the path from the grid connection point to the target power grid node (unit: pu), which is obtained by cumulative calculation based on line resistance and reactance parameters and reflects the electrical distance.
[0036] In the embodiments of the present invention, through the topological path tracing algorithm, starting from the grid connection point of the virtual power plant, the distribution network is traversed level by level to the target power grid node, and the impedance values of all line segments on the path (unit: Ω) are accumulated to obtain the cumulative value of line impedances. The specific process is as follows: Call the power grid geographic information system database to obtain line resistance and reactance parameters, and perform arithmetic accumulation on the modulus values of the impedances of each line on the path.
[0037] Step 202: Extract the target impedance components in the cumulative value of line impedances that account for more than a set threshold, and use the target impedance components as the dominant impedance components; In this step, the set threshold refers to a preset impedance contribution threshold (typical value 15%), which is used to screen key lines. The target impedance component refers to the impedance value of a single line on the path, which needs to satisfy that its proportion in the cumulative value exceeds the set threshold, reflecting the component that has a significant impact on the total impedance. The dominant impedance component refers to the vector sum of the target impedance components (unit: Ω), which characterizes the equivalent impedance of the key line.
[0038] In the embodiments of the present invention, the contribution analysis method is used to process the cumulative value of impedances. Specifically, divide the impedance of each line in the line by the cumulative value of line impedances to obtain the proportion. Assume a set threshold, usually 15%, and screen out the line impedance components whose proportion exceeds this set threshold. For example, the impedance of line L8 accounts for 18%. Extract these components to form the target impedance components, and take their vector sum to generate the dominant impedance components (unit: Ω).
[0039] Step 203: Based on the dominant impedance components and the short-circuit capacity data of the grid connection point, calculate the basic voltage sensitivity factors of each power grid node in the access area of the virtual power plant; In this step, the short-circuit capacity data refers to the apparent power (unit: MVA) at the point of common coupling during a three-phase short circuit, which is calculated based on power system transient stability analysis software. The basic voltage sensitivity factor refers to the preliminary calculated voltage-reactive power sensitivity reference (dimensionless).
[0040] In an embodiment of the present invention, the conductor material property parameters, measured ambient temperature value, spatial length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component are obtained to calculate the reference resistivity value, temperature drift compensation coefficient, and line geometric characteristic value; the line geometric characteristic value and the temperature drift compensation coefficient are multiplied to obtain the physical correction factor; according to the dominant impedance component and the short-circuit capacity data of the point of common coupling, the electrical characteristic factor is calculated, and combined with the physical correction factor, the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant is calculated.
[0041] Step 204: Calculate the weight factor according to the real-time current-carrying capacity and rated current-carrying parameters of the line associated with the dominant impedance component; In this step, the real-time current-carrying capacity refers to the effective value of the current measured by the line current transformer (unit: A), which is updated every 5 seconds and reflects the instantaneous load state of the line. The rated current-carrying parameter refers to the maximum continuous current-carrying capacity allowed by the line design (unit: A), which is determined during the power grid planning stage according to the conductor material and heat dissipation conditions. The weight factor refers to the coefficient (0.5 - 1.5) for correcting the basic sensitivity, which is calculated by comprehensively considering the real-time current-carrying ratio, fouling thermal resistance increment coefficient, heat dissipation efficiency, and harmonic effect, and is used to quantify the line safety margin.
[0042] In an embodiment of the present invention, the real-time current-carrying capacity, rated current-carrying parameters, and historical maximum current-carrying record of the line associated with the dominant impedance component are obtained, the instantaneous approximation ratio of the real-time current-carrying capacity to the rated current-carrying parameter is calculated, and at the same time, based on the historical maximum current-carrying record, the current-carrying fluctuation coefficient is determined; the conductor surface fouling grade and ambient temperature and humidity parameters of the line associated with the dominant impedance component are obtained to calculate the fouling thermal resistance increment coefficient and air heat dissipation efficiency; according to the dominant impedance component, the skin effect depth correction coefficient is calculated; the instantaneous approximation ratio, current-carrying fluctuation coefficient, fouling thermal resistance increment coefficient, air heat dissipation efficiency, and skin effect depth correction coefficient are multiplied to obtain the comprehensive current-carrying capacity factor, and combined with the dominant impedance reference value, the weight factor is calculated.
[0043] Step 205: Superimpose the basic voltage sensitivity factor and the weight factor to generate the voltage sensitivity parameter of each grid node in the access area of the virtual power plant; In an embodiment of the present invention, the basic voltage sensitivity factor and the weight factor are superimposed and calculated according to the linear weighting formula, and the result is normalized to the standard interval [0, 1], and finally the voltage sensitivity parameter is obtained.
[0044] In the embodiment of the present invention, for the problem that the traditional centralized scheduling method has poor adaptability to the change of the power grid topology, by extracting the dominant impedance component and dynamically correcting the voltage sensitivity in combination with the weight factor, the voltage sensitivity characteristics of the power grid nodes after the access of distributed power sources can be more accurately reflected, the voltage stability control ability of the virtual power plant in a complex power grid environment is improved, and the control deviation caused by model mismatch is reduced.
[0045] The present invention provides a specific embodiment. In step 203, according to the dominant impedance component and the short-circuit capacity data of the grid connection point, calculate the basic voltage sensitivity factors of each power grid node in the access area of the virtual power plant, which specifically includes the following steps: Step 211: Obtain the conductor material property parameters, measured environmental temperature value, spatial length dimension, and conductor cross-sectional area parameters of the line corresponding to the dominant impedance component; In this step, the conductor material property parameters refer to the coded data describing the physical characteristics of the line conductor, including the material type, alloy code number, and resistance temperature coefficient. The measured environmental temperature value refers to the real-time air temperature at the location of the line (unit: °C), which is collected by a platinum resistance temperature sensor and reflects the heat dissipation condition of the conductor. The spatial length dimension refers to the straight-line distance between the starting and ending points of the conductor (unit: km), which is measured by a Leica laser rangefinder and is used to calculate the line resistance. The conductor cross-sectional area parameter refers to the geometric dimension of the conductor cross-section (unit: mm²).
[0046] In the embodiment of the present invention, the line ledger data is called through the power grid asset management system to extract the conductor material property parameters; the measured environmental temperature value (unit: °C) is collected in real time by a temperature and humidity sensor; the spatial length dimension (unit: km) is measured by a laser rangefinder; and the conductor cross-sectional area parameter (unit: mm²) is measured by a caliper.
[0047] Step 212: Calculate the reference resistivity value and the temperature drift compensation coefficient according to the conductor material property parameters and the measured environmental temperature value; In this step, the reference resistivity value refers to the resistance value of the conductor per unit length at the standard temperature (20 °C) (unit: Ω·mm² / m), which is obtained by looking up the material resistivity comparison table. The temperature drift compensation coefficient refers to the resistivity correction multiple caused by temperature change and reflects the influence of temperature rise on the resistance.
[0048] In the embodiment of the present invention, the material resistivity database is queried, and the 20 °C reference resistivity value is obtained by matching the material code. It is calculated according to the formula: temperature drift compensation coefficient α = 1 + β×(T - 20), where β is the material temperature coefficient and T is the measured temperature (unit: °C).
[0049] Step 213: Calculate the line geometric characteristic value according to the spatial length dimension and the conductor cross-sectional area parameter; In this step, the line geometric characteristic value refers to the ratio of length to cross-sectional area (unit: m -1 ), and its physical meaning is the equivalent length per unit cross-sectional area.
[0050] In the embodiment of the present invention, the spatial length dimension is divided by the conductor cross-sectional area parameter to obtain the geometric characteristic value (unit: km / mm²), which is dimensionless but is treated as a numerical value in subsequent calculations.
[0051] Step 214: Perform a multiplication operation on the line geometric characteristic value and the temperature drift compensation coefficient to obtain a physical correction factor; In this step, the physical correction factor refers to a correction coefficient that comprehensively considers the geometric characteristics and temperature effects and is used to compensate for changes in the physical state of the line.
[0052] In the embodiment of the present invention, the geometric characteristic value is multiplied by the temperature drift compensation coefficient to obtain the physical correction factor. For example, when the geometric characteristic value is 11666.67 and the temperature drift compensation coefficient is 1.062, the physical correction factor = 11666.67×1.062 = 12390.0, which is dimensionless.
[0053] Step 215: Calculate the electrical characteristic factor based on the dominant impedance component and the short-circuit capacity data of the grid connection point, and combine the physical correction factor to calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant; In this step, the electrical characteristic factor refers to the ratio of the dominant impedance component to the square root of the short-circuit capacity data, which is a dimensionless parameter characterizing the grid strength.
[0054] In the embodiment of the present invention, the electrical characteristic factor = dominant impedance component÷(square root of the short-circuit capacity data); the basic sensitivity factor = electrical characteristic factor×physical correction factor.
[0055] The embodiment of the present invention solves the sensitivity drift caused by the day-night temperature difference through temperature compensation; geometric correction eliminates the sensitivity attenuation of long-distance lines; the combined calculation of the physical correction factor and the electrical characteristic factor helps to improve the voltage regulation success rate of the subsequent reactive power resource dynamic distribution strategy.
[0056] The present invention provides a specific embodiment. In step 204, a weight factor is calculated according to the real-time current-carrying capacity and the rated current-carrying parameter of the line associated with the dominant impedance component, which specifically includes the following steps: Step 221: Obtain the real-time current-carrying capacity, the rated current-carrying parameter, and the historical maximum current-carrying record of the line associated with the dominant impedance component; In this step, the historical maximum current-carrying record refers to the peak current (unit: A) of the line within the statistical period (usually 30 days). The maximum value is extracted from the current-carrying data stored every 5 minutes by the data acquisition and monitoring system, reflecting the extreme load capacity of the line.
[0057] In the embodiment of the present invention, the current-carrying capacity (unit: A) is collected in real time through the current transformer on the line, the rated current-carrying parameters (unit: A) are obtained by reading the grid equipment ledger, and the maximum current-carrying record (unit: A) in the past 30 days is retrieved from the historical database of the data acquisition and monitoring control system.
[0058] Step 222: Calculate the instantaneous approximation ratio of the real-time current-carrying capacity to the rated current-carrying parameters, and at the same time, based on the historical maximum current-carrying record, determine the current-carrying fluctuation coefficient; In this step, the instantaneous approximation ratio refers to the ratio of the real-time current-carrying capacity to the rated current-carrying parameters, which is a dimensionless quantity (in the range of 0 to 1.2) and is used to quantify the current load rate. The current-carrying fluctuation coefficient refers to the ratio of the historical maximum current-carrying capacity to the rated current-carrying capacity multiplied by the fluctuation weight, reflecting the probability of the line overload risk.
[0059] In the embodiment of the present invention, the instantaneous approximation ratio = real-time current-carrying capacity ÷ rated current-carrying parameters; the current-carrying fluctuation coefficient = historical maximum current-carrying record ÷ rated current-carrying parameters × fluctuation weight, where the fluctuation weight is an empirical coefficient with a value of 0.7.
[0060] Step 223: Obtain the conductor surface pollution level and ambient temperature and humidity parameters of the line associated with the dominant impedance component; In this step, the conductor surface pollution level is determined according to the measured values of equivalent salt deposit density and ash deposit density, including levels 0-4, corresponding to different severities of pollution. The ambient temperature and humidity parameters refer to the real-time temperature (unit: °C) and relative humidity (unit: %) at the location where the line is located, which are collected through the wireless sensor network and used to calculate the heat dissipation conditions.
[0061] In the embodiment of the present invention, the conductor surface pollution level is identified by the ultraviolet imager carried by the unmanned aerial vehicle. For example, when the salt deposit density is 0.2 mg / cm², the conductor surface pollution level is 3; the ambient temperature (unit: °C) and relative humidity (unit: %) are collected by the temperature and humidity sensor to obtain the ambient temperature and humidity parameters.
[0062] Step 224: Determine the pollution thermal resistance increment coefficient according to the conductor surface pollution level, and calculate the air heat dissipation efficiency in combination with the ambient temperature and humidity parameters; In this step, the pollution thermal resistance increment coefficient refers to the multiple of the additional thermal resistance caused by pollution, reflecting the impact of pollution accumulation on the temperature rise. The air heat dissipation efficiency refers to the quantified value of the air cooling efficiency, and its physical meaning is the heat dissipation capacity per unit temperature rise.
[0063] In an embodiment of the present invention, a mapping table of the conductor surface pollution level is queried to determine the pollution thermal resistance increment coefficient. For example, the pollution thermal resistance increment coefficient corresponding to the conductor surface pollution level 3 is 1.25; the air heat dissipation efficiency = the reciprocal of the temperature ÷ the humidity compensation term, where the humidity compensation term = 1 - 0.005×H, and H is the relative humidity percentage; the air heat dissipation efficiency = 1÷(T + A)×the humidity compensation term, and A is a constant, which is a value added to achieve temperature scale conversion.
[0064] Step 225: Calculate the leading impedance reference value according to the leading impedance component, and obtain the high-frequency harmonic distortion rate of the line associated with the leading impedance component, so as to calculate the skin effect depth correction coefficient; In this step, the leading impedance reference value refers to the dimensionless value obtained by normalizing the leading impedance component by the system reference impedance, which is used to unify the dimension for calculation. The high-frequency harmonic distortion rate refers to the percentage of the effective value of the harmonic current of the 13th and higher harmonics to the fundamental wave current, which is obtained by fast Fourier transform analysis. The skin effect depth correction coefficient refers to the compensation coefficient for the reduction of the effective cross-sectional area of the conductor caused by high-frequency current, which is used to correct the AC resistance.
[0065] In an embodiment of the present invention, the current signal is analyzed by a power quality analyzer through fast Fourier transform, and the high-frequency harmonic distortion rate is calculated according to the square root of the sum of the squares of the effective values of the 13th to 50th harmonics and the ratio of the fundamental wave current; the skin effect correction coefficient = 1 + 0.02×the high-frequency harmonic distortion rate.
[0066] Step 226: Perform a product operation on the instantaneous approximation ratio, the current-carrying fluctuation coefficient, the pollution thermal resistance increment coefficient, the air heat dissipation efficiency, and the skin effect depth correction coefficient to obtain the comprehensive current-carrying capacity factor, and combine the leading impedance reference value to calculate the weight factor; In this step, the comprehensive current-carrying capacity factor refers to a composite coefficient that integrates the influences of load, pollution, heat dissipation, and harmonics, which is dimensionless and characterizes the line safety margin.
[0067] In an embodiment of the present invention, the comprehensive current-carrying capacity factor = the instantaneous approximation ratio × the current-carrying fluctuation coefficient × the pollution thermal resistance increment coefficient × the air heat dissipation efficiency × the skin effect depth correction coefficient, and the weight factor = the comprehensive current-carrying capacity factor ÷ the leading impedance reference value.
[0068] In an embodiment of the present invention, by dynamically associating the physical state of the line (such as pollution level, temperature and humidity) with electrical parameters (such as current-carrying capacity fluctuation), multi-dimensional dynamic input is provided for the weight factor, enabling the reactive power resource dynamic allocation strategy to respond to grid load changes and environmental disturbances in real time, and improving the adaptability of the virtual power plant to the volatility of distributed power sources.
[0069] The present invention provides a specific embodiment. In step 103, according to the real-time impedance spectrum measurement data, locate the weak line section associated with the power grid node, which specifically includes the following steps: Step 301: Analyze the real-time impedance spectrum measurement data to obtain the impedance amplitude and the phase angle change amount; In this step, the impedance amplitude refers to the modulus value of the line impedance at the fundamental frequency (unit: Ω), which reflects the electrical strength of the line. The phase angle change amount refers to the difference between the current phase angle and the reference phase angle (unit: °). The reference value is the measurement value at the initial stage of line operation and is used to diagnose loose connections. For example, >5° indicates a risk.
[0070] In the embodiment of the present invention, a digital signal processor is used to execute an operation process including frequency-domain data extraction, complex impedance decomposition, impedance amplitude calculation, and phase angle change amount calculation. Specifically, first, perform a 1024-point fast Fourier transform on the real-time impedance spectrum measurement data in the 0.1 - 10 kHz frequency band to locate the 50 Hz fundamental frequency point; extract the complex impedance at the fundamental frequency point: real part = the real part value of the spectrum; imaginary part = the imaginary part value of the spectrum; impedance amplitude = square root (square of the real part + square of the imaginary part); current phase angle = arctangent function (imaginary part / real part) × (180 / π); phase angle change amount = current phase angle - historical reference phase angle.
[0071] Step 302: Calculate the fluctuation coefficient of the phase angle change amount within a preset time window, and at the same time calculate the deviation amount between the impedance amplitude and the historical reference amplitude. Mark the corresponding line segment with a deviation amount exceeding the impedance anomaly threshold and the fluctuation coefficient exceeding the preset stability threshold as a candidate weak section; In this step, the fluctuation coefficient refers to the absolute ratio of the standard deviation to the mean value of the phase angle change amount within a preset time window. >2.0 indicates mechanical instability. The historical reference amplitude refers to the 30-day moving average value of the impedance amplitude in the healthy state of the line (unit: Ω), which is calculated based on the historical database of the data acquisition and monitoring control system. The impedance anomaly threshold refers to the preset impedance deviation rate threshold (usually 20%), and the calculation formula is: |current amplitude - historical reference| / historical reference × 100%. The preset stability threshold refers to the safety limit value of the phase angle fluctuation coefficient (usually 2.0). Exceeding this value is determined as structural instability. The candidate weak section refers to the line segment that meets the double-threshold conditions, and the data structure is [line ID, start point, end point], with a length of 0.5 - 2 km.
[0072] In the embodiment of the present invention, within a 15-minute time window, calculate the fluctuation coefficient, and the fluctuation coefficient = |standard deviation of the phase angle change amount ÷ mean value of the phase angle change amount|; deviation amount = (current impedance amplitude - historical reference amplitude) ÷ historical reference amplitude; when the deviation rate > impedance anomaly threshold and the fluctuation coefficient > preset stability threshold, mark the corresponding line segment as a candidate weak section.
[0073] Step 303: Encode the candidate weak section as a directed line segment according to the starting power grid node and the ending power grid node of the candidate weak section; In this step, the starting power grid node and the ending power grid node refer to the topological endpoints of the line segment, and the node numbering rule is voltage level - region - serial number. A directed line segment refers to a form of representation of a line segment with a direction, defined as an ordered node pair [starting point, ending point].
[0074] In the embodiment of the present invention, according to the distribution network topology database, the starting power grid node (such as NodeA) and the ending power grid node (such as NodeB) of the candidate weak section are combined into an ordered pair to generate a directed line segment, and the specific data structure is [line segment ID: L15, starting point: NodeA, ending point: NodeB].
[0075] Step 304: Traverse all directed line segments. If the ending node of the first directed line segment is the same as the starting node of the second directed line segment, then merge the first directed line segment and the second directed line segment into a target line segment, and combine all target line segments to generate a weak line section; In this step, the first directed line segment and the second directed line segment refer to adjacent line segment objects in the merging operation, and the ending point of the first directed line segment must be equal to the starting point of the second directed line segment. The target line segment refers to the continuous line segment formed after merging, containing ≥2 nodes.
[0076] In the embodiment of the present invention, traverse all the first directed line segments. If the first directed line segment (such as L15 ending point = NodeB) and the second directed line segment (such as L16 starting point = NodeB) have the same nodes; merge them into a continuous target line segment [NodeA, NodeB, NodeC]; recursively execute until there are no line segments with overlapping nodes, and output the final weak line section.
[0077] The embodiment of the present invention solves the pain point that it is difficult for traditional methods to quickly identify key risk points of the power grid through real - time impedance spectrum data analysis and weak line section positioning, and realizes the accurate positioning of the vulnerable areas of the power grid; provides a clear optimization direction for the subsequent dynamic allocation of reactive power resources, and enhances the emergency response ability of the virtual power plant under extreme working conditions.
[0078] The present invention provides a specific embodiment. Step 104, according to the voltage sensitivity parameter and the information of the weak line section, use a deep reinforcement learning agent to generate a dynamic reactive power resource allocation strategy, which specifically includes the following steps: Step 401: Create a feature group for each power grid node. The feature group includes a node optimization sensitivity parameter and a weak node, and the weak node is determined according to whether the corresponding node is the starting power grid node or the ending power grid node of the weak line section; In this step, the feature group refers to a vector container that describes the electrical state of grid nodes, with a data structure of {node ID, node optimization sensitivity parameter, weak node flag}, where the node optimization sensitivity parameter is obtained by correcting with a risk enhancement factor. The node optimization sensitivity parameter refers to the sensitivity value (dimensionless) enhanced by voltage offset integration and topological weight, and the calculation formula is: optimization sensitivity = base sensitivity × (1 + voltage offset integration × connection weight), and the numerical range is 0.05 - 0.15. A weak node refers to a grid node corresponding to the endpoint of a weak line section.
[0079] In the embodiment of the present invention, according to the node connection relationship in the weak line section, the connection weight coefficient between the starting grid node and the ending grid node is calculated; the connection weight coefficient corresponding to the grid node marked as a weak node is multiplied by the voltage offset integral value to generate a risk enhancement factor, and combined with the voltage sensitivity parameter, the node optimization sensitivity parameter is calculated, and it is combined with the weak node to form the core element of the feature group, and the core element of the feature group is combined with the hierarchical position coding of the grid node in the virtual power plant to generate the feature group of each grid node.
[0080] Step 402: Determine the target grid node connected to each reactive power resource device according to a preset device-node association table; In this step, the preset device-node association table refers to a predefined device topology relationship database that stores the mapping between the reactive power resource device ID and its physical connection node. The target grid node refers to the distribution network node number directly accessed by the reactive power resource device, which is obtained by querying the device-node association table.
[0081] In the embodiment of the present invention, the preset device-node association table is read, the mapping relationship between the reactive power resource device ID and the grid node is parsed, and the target grid node connected to each reactive power resource device is determined.
[0082] Step 403: Use a pre-trained deep reinforcement learning agent to analyze the feature group and the target grid node to generate an adjustment instruction value for each reactive power resource device; In this step, the adjustment instruction value refers to the normalized control instruction value (in the range of [-1, 1]) output by the deep reinforcement learning agent, where a positive number indicates emitting reactive power / discharging, and a negative number indicates absorbing reactive power / charging.
[0083] In the embodiment of the present invention, the feature group is sorted into a 128-dimensional vector according to the node number, and the target node information is encoded as a 32-dimensional one-hot vector; a 160-dimensional state vector is input through the pre-trained deep reinforcement learning agent, and an adjustment instruction value in the continuous action space [-1, 1] is output, where the action output layer uses a hyperbolic tangent activation function to ensure that the instruction value is in the range of [-1, 1].
[0084] Step 404: Convert the adjustment command value into a reactive power setting value of the distributed power source and a charge-discharge power setting value of the energy storage converter, so as to generate a dynamic reactive resource allocation strategy according to the reactive power setting value and the charge-discharge power setting value; In this step, the reactive power setting value refers to the actual control parameter of the distributed power source (unit: kVar), which is calculated by multiplying the adjustment command value by the rated capacity, and the range is [-rated value, +rated value]. The charge-discharge power setting value refers to the active power command of the energy storage converter (unit: kW), which is obtained by multiplying the adjustment command value by the rated power, and a negative value indicates charging.
[0085] In the embodiment of the present invention, the reactive power setting value of the distributed power source = adjustment command value × equipment rated capacity; the charge-discharge power setting value of the energy storage converter = adjustment command value × rated power; finally, the two setting values are encapsulated into a dynamic reactive resource allocation strategy.
[0086] The embodiment of the present invention generates a dynamic reactive resource allocation strategy through a deep reinforcement learning agent, breaking through the limitations of traditional heuristic algorithms in real-time performance and self-adaptability, improving the real-time performance and coordination of the multi-objective optimal scheduling of the virtual power plant, and meeting the dynamic balance requirements after the access of a high proportion of renewable energy.
[0087] The present invention provides a specific embodiment. In step 401, a feature group is created for each grid node. The feature group includes a node optimization sensitivity parameter and a weak node. The weak node is determined according to whether the corresponding node is the starting grid node or the ending grid node of a weak line section. The specific steps are as follows: Step 411: Collect the voltage fluctuation trajectory data of the grid node within a preset time window to calculate the voltage offset integral value; In this step, the voltage fluctuation trajectory data refers to the continuous sampling sequence (unit: pu) of the grid node voltage within a preset time window (15 minutes), which is collected by a phasor measurement device at a rate of ≥30 frames / second and reflects the voltage transient fluctuation characteristics. The voltage offset integral value refers to the cumulative amount of the voltage deviating from the reference value (unit: pu·s), which quantifies the severity of voltage over-limit.
[0088] In the embodiment of the present invention, the voltage fluctuation trajectory data (unit: pu) of the grid node within a 15-minute preset time window is collected by a phasor measurement device at a rate of 50 frames per second. The trapezoidal integration method is used to calculate the voltage offset integral value, and the voltage offset integral value = ∑[|actual voltage - reference voltage| × time interval].
[0089] Step 412: Calculate the connection weight coefficient between the starting grid node and the ending grid node according to the node connection relationship in the weak line section; In this step, the node connection relationship refers to the topological connection structure of the power grid nodes in the weak line section, which is represented as an adjacency matrix and generated based on graph theory algorithms. The connection weight coefficient refers to the importance index (0-1) of the node in the weak section, reflecting the risk of fault propagation.
[0090] In the embodiment of the present invention, the node adjacency matrix is extracted from the topology of the weak line section, and the degree centrality algorithm is used to calculate the connection weight coefficient. The connection weight coefficient = (the number of lines connected to the current node) / (the total number of lines in the weak section).
[0091] Step 413: Multiply the connection weight coefficient corresponding to the power grid node marked as a weak node by the voltage offset integral value to generate a risk enhancement factor, and combine the voltage sensitivity parameter to calculate the node optimization sensitivity parameter; In this step, the risk enhancement factor refers to the risk gain coefficient (≥0) only for weak nodes, which is obtained by multiplying the connection weight coefficient by the voltage offset integral value and is used to enhance the sensitivity weight of high risks.
[0092] In the embodiment of the present invention, the risk enhancement factor is calculated only for weak nodes (i.e., marked = 1); the risk enhancement factor = connection weight coefficient × voltage offset integral value; the node optimization sensitivity parameter = voltage sensitivity parameter × (1 + risk enhancement factor), and non-weak nodes remain unchanged.
[0093] Step 414: Combine the node optimization sensitivity parameter and the weak node into the core elements of the feature group, and combine the core elements of the feature group with the hierarchical position coding of the power grid node in the virtual power plant to generate the feature group of each power grid node; In this step, the core elements of the feature group refer to the core data units describing the electrical state of the node, and the data structure is a binary tuple [node optimization sensitivity parameter, weak node]. The hierarchical position coding refers to the hierarchical identifier of the node in the control architecture of the virtual power plant, which is used to distinguish the control priorities.
[0094] In the embodiment of the present invention, the core elements of the feature group are encapsulated as a binary tuple [node optimization sensitivity parameter, weak node]; according to the three-level architecture of the power grid node in the virtual power plant, the coding is assigned. Among them, the coding of the main network access layer is 100, the coding of the distribution network hub layer is 200, and the coding of the user access layer is 300; finally, the core elements of the feature group and the hierarchical coding are combined into a structured feature group [sensitivity value, weak node, hierarchical coding].
[0095] The embodiments of the present invention solve the problem of insufficient characterization of the key characteristics of power grid nodes by traditional scheduling methods; the deep reinforcement learning agent can more comprehensively capture the topological characteristics and operating status of the power grid, generate more accurate reactive power resource adjustment instruction values; improve the adaptive scheduling ability of the virtual power plant in a complex power grid environment, and reduce the system stability risk caused by the volatility of distributed power sources.
[0096] Figure 2 FIG. is a schematic structural diagram of a virtual power plant multi-objective optimal scheduling system integrating deep reinforcement learning provided by an embodiment of the present invention. As Figure 2 shown, the system includes: An acquisition module 21, configured to acquire short-circuit capacity data and real-time impedance spectrum measurement data of the connection point of the virtual power plant. A calculation module 22, configured to calculate voltage sensitivity parameters of each power grid node in the access area of the virtual power plant based on the short-circuit capacity data. A positioning module 23, configured to locate a weak line section associated with the power grid node according to the real-time impedance spectrum measurement data. A generation module 24, configured to generate a dynamic reactive power resource allocation strategy by using a deep reinforcement learning agent according to the voltage sensitivity parameters and the information of the weak line section, where the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss.
[0097] A control module 25, configured to coordinately control the reactive power output power of the distributed power source and the charge and discharge power of the energy storage converter according to the dynamic reactive power resource allocation strategy to achieve multi-objective optimal scheduling of the virtual power plant.
[0098] Figure 2 The virtual power plant multi-objective optimal scheduling system integrating deep reinforcement learning can execute Figure 1 the virtual power plant multi-objective optimal scheduling method described in the embodiment shown, and its implementation principle and technical effects will not be elaborated. For the virtual power plant multi-objective optimal scheduling system integrating deep reinforcement learning in the above embodiment, the specific ways for each module and unit to execute operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0099] In a possible design, Figure 2 the virtual power plant multi-objective optimal scheduling system described in the embodiment shown can be implemented as a computing device. As Figure 3 shown, the computing device may include a storage component 31 and a processing component 32; The storage component 31 stores one or more computer instructions, where the one or more computer instructions are called and executed by the processing component 32.
[0100] The processing component 32 is used for the above Figure 1 A virtual power plant multi-objective optimal scheduling method integrating deep reinforcement learning in the above-described embodiment.
[0101] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0102] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0103] Of course, the computing device may also necessarily include other components, such as input / output interfaces, display components, communication components, etc.
[0104] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc.
[0105] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0106] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above processing component, storage component, etc. may be basic server resources leased or purchased from the cloud computing platform.
[0107] An embodiment of the present invention also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 A virtual power plant multi-objective optimal scheduling method integrating deep reinforcement learning in the above-described embodiment.
[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-objective optimal scheduling method for virtual power plants integrating deep reinforcement learning, characterized in that Including: Obtain the short - circuit capacity data and real - time impedance spectrum measurement data of the connection point of the virtual power plant; Based on the short - circuit capacity data, calculate the voltage sensitivity parameters of each grid node in the access area of the virtual power plant; According to the real - time impedance spectrum measurement data, locate the weak line sections associated with the grid nodes; According to the voltage sensitivity parameters and the information of the weak line sections, use a deep reinforcement learning agent to generate a dynamic reactive power resource allocation strategy, where the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss; According to the dynamic reactive power resource allocation strategy, coordinately control the reactive power output power of distributed power sources and the charge - discharge power of energy storage converters to achieve multi - objective optimal scheduling of the virtual power plant.
2. The method according to claim 1, characterized in that Based on the short - circuit capacity data, calculating the voltage sensitivity parameters of each grid node in the access area of the virtual power plant includes: Calculate the cumulative value of the line impedance from the connection point of the virtual power plant to each grid node in the access area of the virtual power plant; Extract the target impedance component in the cumulative line impedance value whose proportion exceeds the set threshold, and use the target impedance component as the dominant impedance component; According to the dominant impedance component and the short - circuit capacity data of the connection point, calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant; Calculate the weight factor according to the real - time current - carrying capacity and rated current - carrying parameters of the line associated with the dominant impedance component; Superimpose the basic voltage sensitivity factor and the weight factor to generate the voltage sensitivity parameters of each grid node in the access area of the virtual power plant.
3. The method according to claim 2, wherein According to the dominant impedance component and the short - circuit capacity data of the connection point, calculating the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant includes: Obtain the conductor material property parameters, measured environmental temperature value, spatial length dimension, and conductor cross - sectional area parameters of the line corresponding to the dominant impedance component; According to the conductor material property parameters and the measured environmental temperature value, calculate the reference resistivity value and the temperature drift compensation coefficient; According to the spatial length dimension and the conductor cross - sectional area parameters, calculate the line geometric characteristic value; Perform a product operation on the line geometric characteristic value and the temperature drift compensation coefficient to obtain the physical correction factor; According to the dominant impedance component and the short - circuit capacity data of the connection point, calculate the electrical characteristic factor, and combine the physical correction factor to calculate the basic voltage sensitivity factor of each grid node in the access area of the virtual power plant.
4. The method according to claim 2, wherein According to the real - time current - carrying capacity and rated current - carrying parameters of the line associated with the dominant impedance component, calculating the weight factor includes: Obtain the real - time current - carrying capacity, rated current - carrying parameters, and historical maximum current - carrying record of the line associated with the dominant impedance component; Calculate the instantaneous approximation ratio of the real - time current - carrying capacity to the rated current - carrying parameter, and at the same time, based on the historical maximum current - carrying record, determine the current - carrying fluctuation coefficient; Obtain the conductor surface pollution level and environmental temperature and humidity parameters of the line associated with the dominant impedance component; According to the conductor surface pollution level, determine the pollution thermal resistance increment coefficient, and calculate the air heat dissipation efficiency in combination with the environmental temperature and humidity parameters; Based on the dominant impedance component, a dominant impedance reference value is calculated, and the high-frequency harmonic distortion rate of the line associated with the dominant impedance component is obtained to calculate the skin effect depth correction factor; The instantaneous approximation ratio, the current-carrying fluctuation coefficient, the pollution thermal resistance increment coefficient, the air heat dissipation efficiency, and the skin effect depth correction factor are multiplied to obtain a comprehensive current-carrying capacity factor, and in combination with the dominant impedance reference value, a weight factor is calculated.
5. The method according to claim 1, characterized in that Based on the real-time impedance spectrum measurement data, the weak line sections associated with the power grid nodes are located, including: Analyze the real-time impedance spectrum measurement data to obtain the impedance amplitude and the phase angle change amount; Calculate the fluctuation coefficient of the phase angle change amount within a preset time window, and at the same time calculate the deviation amount between the impedance amplitude and the historical reference amplitude. Mark the corresponding line section where the deviation amount exceeds the impedance anomaly threshold and the fluctuation coefficient exceeds the preset stability threshold as a candidate weak section; According to the starting power grid node and the ending power grid node of the candidate weak section, encode the candidate weak section as a directed line segment; Traverse all directed line segments. If the ending node of the first directed line segment is the same as the starting node of the second directed line segment, then merge the first directed line segment and the second directed line segment into a target line segment, and combine all target line segments to generate a weak line section.
6. The method according to claim 1, wherein Based on the voltage sensitivity parameter and the information of the weak line section, use a deep reinforcement learning agent to generate a reactive power resource dynamic allocation strategy, including: Create a feature group for each power grid node, where the feature group includes a node optimization sensitivity parameter and a weak node, and the weak node is determined according to whether the corresponding node is the starting power grid node or the ending power grid node of the weak line section; According to a preset device node association table, determine the target power grid node connected to each reactive power resource device; Use a pre-trained deep reinforcement learning agent to analyze the feature group and the target power grid node to generate an adjustment command value for each reactive power resource device; Convert the adjustment command value into a reactive power setting value of the distributed power source and a charge and discharge power setting value of the energy storage converter, so as to generate a reactive power resource dynamic allocation strategy according to the reactive power setting value and the charge and discharge power setting value.
7. The method according to claim 6, wherein Create a feature group for each power grid node, where the feature group includes a node optimization sensitivity parameter and a weak node, and the weak node is determined according to whether the corresponding node is the starting power grid node or the ending power grid node of the weak line section, including: Collect the voltage fluctuation trajectory data of the power grid node within a preset time window to calculate the voltage offset integral value; According to the node connection relationship in the weak line section, calculate the connection weight coefficient between the starting power grid node and the ending power grid node; Multiply the connection weight coefficient corresponding to the power grid node marked as a weak node by the voltage offset integral value to generate a risk reinforcement factor, and in combination with the voltage sensitivity parameter, calculate the node optimization sensitivity parameter; Combining the optimized sensitivity parameter of the node and the weak node into the core element of the feature group, and combining the core element of the feature group with the hierarchical position coding of the power grid node in the virtual power plant to generate the feature group of each power grid node.
8. A virtual power plant multi-objective optimal scheduling system integrating deep reinforcement learning, characterized in that Including: An acquisition module for acquiring the short-circuit capacity data and real-time impedance spectrum measurement data of the grid connection point of the virtual power plant; A calculation module for calculating the voltage sensitivity parameters of each power grid node in the access area of the virtual power plant based on the short-circuit capacity data; A positioning module for positioning the weak line section associated with the power grid node according to the real-time impedance spectrum measurement data; A generation module for generating a reactive power resource dynamic allocation strategy by using a deep reinforcement learning agent according to the voltage sensitivity parameter and the information of the weak line section, where the deep reinforcement learning agent aims to maximize the voltage regulation compensation benefit and minimize the device regulation loss; A control module for coordinately controlling the reactive power output power of the distributed power source and the charge and discharge power of the energy storage converter according to the reactive power resource dynamic allocation strategy to achieve the multi-objective optimal scheduling of the virtual power plant.
9. A computing device, characterized in that, Including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-objective optimal scheduling method of a virtual power plant integrating deep reinforcement learning as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, A computer program is stored, and when the computer program is executed by the computer, it implements a multi-objective optimal scheduling method of a virtual power plant integrating deep reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual power plant dynamic regulation and control method based on deep reinforcement learning
CN119543316A
Power system multi-target scheduling optimization control method and system based on reinforcement learning
CN119695928A
Distributed energy storage equipment predictive maintenance system and method based on AI
CN119919125A
Multi-target scheduling optimization method and system for virtual power plant to participate in electricity market
CN120087801A
Power grid reactive voltage control method based on two-stage deep reinforcement learning
US20210356923A1
Cited By
Energy storage system scheduling and operation optimization method under virtual power plant platform
CN120601428A
Distributed energy intelligent optimization scheduling system and method based on virtual power plant
CN120767941A
Data analysis method in virtual power plant
CN120931048A
Virtual power plant distributed cooperative control method and system based on multi-agent game and impedance perception
CN121367330A
Virtual power plant distributed resource cooperative control method and device oriented to power distribution network
CN121485157A