A multi-source data fusion distributed photovoltaic capacity online identification method and system
Patent Information
- Application Number
- CN202610754500.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明旨在解决现有技术中分布式光伏容量辨识数据来源单一、实时性差、可靠性不足的技术问题,提供一种多源数据融合的分布式光伏容量在线辨识方法及系统,无需增设专用量测设备,仅利用配电网现有量测系统、公共气象服务及电网拓扑信息,即可实现分布式光伏安装容量的在线准确辨识
1、无需专用量测设备,降低系统改造成本:现有分布式光伏容量辨识方法多依赖智能电表、微型同步相量测量装置等专用量测设备的加装,存在投资大、施工难、用户配合度低等问题。本发明仅利用配电网现有调度控制系统量测数据、公共气象服务数据及电网GIS拓扑数据,通过多源信息融合与智能算法实现容量辨识,无需对现场设备进行任何硬件改造,显著降低了技术推广门槛和工程实施成本。
Smart Images

Figure CN122823747A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system operation and control technology, specifically relating to a method and system for online identification of distributed photovoltaic capacity through multi-source data fusion. Background Technology
[0002] Distributed photovoltaic (PV) systems are characterized by their large number, dispersed layout, and highly random power output. Their "black box" operation poses a severe challenge to the safe and stable operation of the distribution network. Accurately grasping the installed capacity information of distributed PV systems in the distribution network is an important foundation for conducting distributed PV integration capacity assessments, distribution network voltage over-limit early warnings, distributed energy dispatch planning, and distribution network planning scheme verification.
[0003] Currently, the acquisition of distributed photovoltaic (PV) capacity information mainly relies on the following methods: First, statistics based on user installation application files, but this has problems such as discrepancies between the applied capacity and the actual installed capacity, and lagging file updates; Second, direct collection of PV inverter data through dedicated measurement devices (such as smart meters and micro synchronous phasor measurement devices), but this is difficult to promote on a large scale due to factors such as high retrofit costs, poor communication conditions, and user privacy concerns; Third, indirect identification based on state estimation or data-driven methods, but existing methods often only use a single data source (such as only measurement data or only meteorological data), without fully considering the spatiotemporal complementarity of multi-source data, resulting in low identification accuracy and poor adaptability.
[0004] Furthermore, most existing identification methods employ offline batch processing, which cannot adapt to the demands of scenarios involving frequent commissioning and decommissioning of distributed photovoltaic (PV) power and rapidly changing weather conditions. While some methods possess online capabilities, they lack effective data quality verification and result feedback mechanisms, making it difficult to guarantee the reliability of the identification results. Therefore, there is an urgent need for a distributed PV capacity identification technology solution that integrates multi-source data, possesses online real-time identification capabilities, and includes a closed-loop verification mechanism. Summary of the Invention
[0005] This invention aims to solve the technical problems of single data source, poor real-time performance, and insufficient reliability in the existing distributed photovoltaic capacity identification technology. It provides a method and system for online identification of distributed photovoltaic capacity by multi-source data fusion. It does not require the addition of dedicated measurement equipment. It can achieve accurate online identification of distributed photovoltaic installation capacity by utilizing only the existing measurement system of the distribution network, public meteorological services, and power grid topology information.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for online identification of distributed photovoltaic capacity through multi-source data fusion includes the following steps: S1. Multi-source data acquisition and preprocessing: Acquire measurement data, meteorological data, and topology data of the target distribution network area; identify and repair bad data in the measurement data; perform spatiotemporal alignment and standardization processing on the meteorological data; verify node correlation in the topology data; and construct a multi-source fusion dataset with a unified spatiotemporal benchmark. The meteorological data includes irradiance data retrieved from satellite cloud images, measured data from ground meteorological stations, and numerical weather prediction data. The spatiotemporal alignment includes mapping meteorological data with different spatial resolutions to distribution network nodes using inverse distance weighting interpolation, and aligning meteorological data with different temporal resolutions to the measurement data timestamps using linear interpolation. S2. Operational status feature extraction and partition aggregation: Based on the multi-source fusion dataset, extract the net load power curve features, voltage sensitivity features and meteorological response features of each node, perform adaptive partitioning according to the electrical distance and feature similarity between nodes, divide the target distribution network area into several identification sub-regions, and aggregate in each sub-region to obtain the equivalent net load sequence and equivalent meteorological sequence; S3. Dynamic Decoupling of Distributed Photovoltaic Output: For each identified sub-region, an equivalent photovoltaic output model considering the coupling relationship of meteorological factors is established. Using the equivalent meteorological sequence and the equivalent net load sequence, the photovoltaic output is separated from the net load through meteorological correlation analysis to generate the photovoltaic output estimation sequence for each sub-region. The separation process uses a sliding time window to dynamically update the model parameters to adapt to the time-varying characteristics of the load. S4. Online Capacity Identification and Iterative Optimization: Based on the photovoltaic power output estimation sequence and the corresponding meteorological irradiance data, a photovoltaic capacity-output mapping relationship is established. A variable step-size gradient descent method is used for iterative solution. Photovoltaic capacity is the optimization variable, and minimizing the fitting residual between the photovoltaic power output estimation sequence and the standard photovoltaic power output curve is the objective function. The step size is determined based on the ratio of the residual change rate of the current iteration step to the residual change rate of the previous iteration step. When the ratio is less than a preset attenuation threshold, the current step size is multiplied by a preset attenuation coefficient. When the residual change directions are opposite in two consecutive iterations, the step size is restored to its initial value, the search direction is adjusted, and smoothing correction is performed through temporal consistency constraints of identification results in adjacent time periods. The online identification results of the distributed photovoltaic capacity of each node are then output. S5. Identification Result Verification and Anomaly Feedback: The online identification result is compared with the preset reasonable capacity range. If the identification result exceeds the reasonable capacity range, the data quality backtracking verification mechanism is triggered, and steps S1 to S4 are re-executed until the identification result meets the convergence condition or reaches the maximum number of iterations. The reasonable capacity range is determined comprehensively based on the rated capacity of the transformers at the distribution network nodes, historical application data, regional solar resource level, and the upper limit of distributed photovoltaic penetration rate.
[0007] Preferably, in step S1, the identification and repair of bad data in the measurement data includes: using a residual detection method based on node power balance to identify abnormal measurement values; for the identified abnormal measurement values, priority is given to using the trend extrapolation value of the historical data of the same node for repair; if historical data is missing, then weighted interpolation of the measurement values of the nodes in the same partition is used for repair, and the weight of the weighted interpolation is determined according to the reciprocal of the electrical distance between nodes.
[0008] Preferably, in step S2, the adaptive partitioning based on the electrical distance and feature similarity between nodes includes: calculating a comprehensive similarity index between nodes, wherein the comprehensive similarity index is composed of the reciprocal of the electrical distance and the cosine similarity of the feature vector, and partitioning is performed using a clustering algorithm based on local density and relative distance, wherein the local density is determined based on the number of neighboring nodes within the electrical distance threshold range of a node, and the relative distance is determined based on the minimum electrical distance from a node to a node with higher local density, so that nodes in the same sub-region have similar electrical coupling characteristics and meteorological response patterns, and the number of sub-regions is automatically determined according to the scale of the distribution network.
[0009] Preferably, in step S4, the smoothing correction through the temporal consistency constraint of the identification results of adjacent time periods includes: establishing a capacity change penalty term for adjacent identification time periods, wherein the penalty term is proportional to the capacity change rate; adding the penalty term to the objective function to form an optimization problem with constraints; and using the exponential smoothing method to post-process the identification results of multiple consecutive time periods to suppress the jump in capacity identification values caused by measurement noise.
[0010] This invention also provides a distributed photovoltaic capacity online identification system based on multi-source data fusion, comprising a data access layer, a data preprocessing module, a feature extraction and partitioning module, an output decoupling module, a capacity identification module, and a result verification module. The data access layer is used to access multi-source data from a distribution network measurement system, a meteorological service system, and a power grid GIS system. The data preprocessing module is used to perform the multi-source data acquisition and preprocessing described in step S1 of claim 1. The feature extraction and partitioning module is used to perform the operating status feature extraction and partitioning aggregation described in step S2 of claim 1. The output decoupling module is used to perform the distributed photovoltaic power output dynamic decoupling described in step S3 of claim 1. The capacity identification module is used to perform the online capacity identification and iterative optimization described in step S4 of claim 1. The result verification module is used to perform the identification result verification and anomaly feedback described in step S5 of claim 1.
[0011] Preferably, the capacity identification module adopts an edge-cloud collaborative architecture, deploying a lightweight identification model on the edge side to perform real-time capacity estimation, and deploying a high-precision identification model on the cloud side to perform periodic verification and model updates. The edge side and the cloud side maintain the consistency of identification parameters through an incremental synchronization mechanism.
[0012] The present invention also provides the application of the above method in the distribution network dispatch and control system, wherein the online identification results are used for distributed photovoltaic acceptance capacity assessment, distribution network voltage over-limit early warning, distributed energy dispatch plan preparation and distribution network planning scheme verification.
[0013] This invention proposes a five-layer progressive online identification mechanism consisting of "multi-source data fusion + feature extraction partitioning + dynamic decoupling of output + capacity iterative identification + result closed-loop verification". Its working principle can be divided into five functional layers: data layer, feature layer, decoupling layer, identification layer and verification layer. Each layer is connected through a standardized data interface to form a complete closed-loop workflow.
[0014] 1. Data Layer: Unified access and governance of multi-source heterogeneous data This invention accesses three types of heterogeneous data sources in parallel through the data access layer: ① Distribution network measurement data: from the dispatch control system (DSCADA), including voltage, active power, and reactive power measurements of each node, with a sampling period of typically 15 minutes; ② Meteorological data: including irradiance retrieved from satellite cloud images (wide spatial coverage but coarse resolution), measured data from ground meteorological stations (high accuracy but sparse stations), and numerical weather forecasts (providing future trends but limited timeliness); ③ Power grid topology data: from the power grid GIS system, including node connection relationships, line impedance parameters, transformer tap positions, and historical installation location information.
[0015] The data preprocessing module performs quality control on the aforementioned multi-source data. For measurement data, a residual detection method based on node power balance is used to identify bad data, prioritizing extrapolation of historical trends from the same node for data repair. When historical data is missing, interpolation based on the inverse of electrical distance weighting is used for repair within the same region. For meteorological data, inverse distance weighting interpolation is used to map data of different spatial resolutions to distribution network nodes, and linear interpolation is used to align data of different temporal resolutions with measurement timestamps, ultimately constructing a multi-source fusion dataset with a unified spatiotemporal benchmark. For topology data, connectivity and electrical consistency checks are performed to eliminate topology anomalies caused by incorrect switch states or parameter distortions.
[0016] 2. Feature Layer: Extraction of runtime status features and adaptive partitioning aggregation This invention decomposes a large-scale distribution network into several identification sub-regions with homogeneous internal characteristics and clear external boundaries, thereby reducing the computational dimensionality of subsequent identification problems.
[0017] The feature extraction and partitioning module is based on a multi-source fusion dataset and extracts node operating status features from three dimensions: ① Net load power curve features: including daily peak-to-valley difference rate, daily average power, power fluctuation coefficient and power ramp rate, reflecting the load behavior pattern of the node; ② Voltage sensitivity features: including voltage sensitivity to active / reactive injection, reflecting the electrical position characteristics of the node in the power grid; ③ Meteorological response features: including the time-delay correlation coefficient between net load and temperature, irradiance and cloud cover index, identifying meteorologically sensitive load components.
[0018] Based on the aforementioned feature vectors, a comprehensive similarity index is calculated between nodes, weighted by the reciprocal of the electrical distance and cosine similarity. Adaptive partitioning is then performed using a clustering algorithm based on local density and relative distance. Local density is determined by the number of neighboring nodes within an electrical distance threshold, while relative distance is determined by the minimum electrical distance from a node to a higher-density node. The number of sub-regions is automatically determined based on the distribution network scale, requiring no manual pre-setting.
[0019] 3. Decoupling layer: Dynamic separation of photovoltaic output and net load. This invention accurately separates the photovoltaic output component from the equivalent net load after polymerization, eliminating the interference of temperature-sensitive loads (such as air conditioning) and radiation-sensitive loads (such as lighting) on the identification results.
[0020] The output decoupling module establishes an equivalent photovoltaic output model for each identified sub-region, taking into account the coupling relationship of meteorological factors. This model uses irradiance, temperature, and cloud cover correction coefficients as inputs and photovoltaic capacity as a linear scaling factor to describe the photovoltaic output characteristics under ideal conditions.
[0021] A multivariate correlation model between net load and meteorological factors was established through meteorological correlation analysis. Partial correlation analysis was used to identify and remove temperature-sensitive and radiation-sensitive load components. The separation process employed a sliding time window to dynamically update model parameters. The window length covered the most recent 30 days of data, and the window moved forward one day after each identification period, ensuring that the model always tracked the seasonal and time-varying characteristics of load features and adapted to scenarios such as differences in air conditioning load between summer and winter.
[0022] 4. Identification Layer: Iterative Optimization Solution of Photovoltaic Capacity This invention uses the decoupled photovoltaic power output estimation sequence and corresponding meteorological irradiance data to invert and solve the equivalent installed capacity of distributed photovoltaic power in each sub-region.
[0023] The capacity identification module establishes a photovoltaic (PV) capacity-output mapping relationship, with the optimization objective of minimizing the fitting residual. It employs a variable-step-size gradient descent method for iterative solution. The optimization variable is PV capacity, and the objective function is the sum of squared residuals between the estimated PV output sequence and the standard PV output curve, plus a time-series consistency penalty term. During iteration, the step size is adaptively adjusted based on the ratio of the residual change rate between the current iteration and the previous iteration: when the ratio is less than a preset decay threshold, the residual descent rate is considered to have slowed down, and the step size is automatically reduced for finer searching; when the residual change directions are opposite in two consecutive iterations, oscillation is identified, the step size is restored to its initial value, and the search direction is reversed. This mechanism effectively avoids the oscillation or divergence phenomena that easily occur in non-convex optimization problems with fixed-step-size gradient descent. Meanwhile, a smoothing correction is performed by constraining the temporal consistency of identification results in adjacent time periods: a penalty term proportional to the rate of capacity change is introduced into the objective function to form an optimization problem with constraints; and an exponential smoothing method is used to post-process the identification results of multiple consecutive time periods, making full use of the physical characteristic that the installed capacity of distributed photovoltaics remains unchanged in the short term and suppressing the jump in identification values caused by measurement noise.
[0024] The capacity identification module adopts an edge-cloud collaborative architecture. A lightweight identification model is deployed at the edge to perform real-time capacity estimation, meeting the requirements of second-level response; a high-precision identification model is deployed in the cloud to perform daily regular verification and model updates; the two sides maintain the consistency of identification parameters through an incremental synchronization mechanism, balancing real-time performance and accuracy.
[0025] 5. Verification Layer: Ensuring Result Reliability and Providing Anomaly Feedback This invention performs engineering rationality verification on the identification results, forming a closed loop of "identification-verification-correction" to ensure that the output results are reliable and usable.
[0026] The result verification module compares the online identification results with a preset reasonable capacity range. The reasonable capacity range is determined comprehensively based on the transformer's rated capacity, historical installation data, regional solar resource levels, and the upper limit of distributed photovoltaic penetration. This prevents absurdly high values due to measurement faults and avoids overlooking actual photovoltaic capacity. If the identification result is within the reasonable range, the final result is directly output along with an identification confidence level indicator. If the identification result exceeds the reasonable range, a data quality backtracking verification mechanism is triggered: the quality of measurement data, the accuracy of meteorological data alignment, and the rationality of adaptive zoning are checked sequentially. After automatically repairing any issues found, the entire process from the data layer to the identification layer is re-executed until the result converges or the maximum number of iterations is reached. If convergence fails after multiple iterations, the node is marked as "identification anomaly," the most recent result is output, and manual intervention is prompted for verification.
[0027] The system displays the online identification results, confidence level, anomaly indicators, and historical trends of distributed photovoltaic capacity at each node through a human-computer interaction interface. The results are then pushed to relevant application modules of the distribution network dispatch and control system to support business scenarios such as distributed photovoltaic capacity assessment, voltage over-limit early warning, dispatch plan preparation, and planning scheme verification.
[0028] This invention features a distinct closed-loop characteristic. This closed-loop mechanism ensures that the system maintains the continuity and reliability of identification results even under disturbances such as measurement anomalies, missing meteorological data, and topology changes, enabling continuous online monitoring of distributed photovoltaic capacity. Compared with existing technologies, this invention has the following advantages: 1. Multi-source data fusion improves identification accuracy: By comprehensively utilizing multi-source meteorological data such as distribution network measurement data, irradiance inversion from satellite cloud images, ground meteorological station measurements, and numerical weather forecasts, a unified dataset is constructed through spatiotemporal alignment and standardization processing. This fully leverages the complementary advantages of different data sources in terms of spatial resolution, temporal resolution, and coverage, effectively overcoming the problems of information loss or insufficient accuracy from single data sources.
[0029] 2. Adaptive partitioning and aggregation reduces computational complexity: Adaptive partitioning is performed based on a comprehensive index of electrical distance and feature similarity, decomposing the large-scale distribution network into several homogeneous sub-regions. Equivalent sequences are aggregated and identified within each sub-region, which not only preserves the electrical coupling characteristics and meteorological response patterns within the region, but also significantly reduces the computational dimension in high-penetration scenarios, meeting the requirements for online real-time performance.
[0030] 3. Improved decoupling accuracy through meteorological-load joint separation: Photovoltaic output is separated from net load through meteorological correlation analysis, and the model parameters are dynamically updated using a sliding time window. This effectively eliminates the interference of temperature-sensitive loads (such as air conditioning loads) and radiation-sensitive loads on the decoupling results, adapts to the seasonal and time-varying characteristics of load characteristics, and improves the accuracy of photovoltaic output estimation.
[0031] 4. Variable step size gradient descent method ensures convergence stability: With the goal of minimizing the fitting residual, the step size is adaptively adjusted according to the ratio of the residual change rate. Combined with the decay threshold judgment and backtracking recovery mechanism, it avoids the oscillation or divergence phenomenon that is prone to occur in the traditional fixed step size gradient descent method in the non-convex optimization problem of capacity identification, and ensures the convergence speed and stability of the iterative solution.
[0032] 5. Temporal consistency constraints suppress the impact of measurement noise: By establishing a capacity change penalty term to form an optimization problem with constraints, and using the exponential smoothing method for post-processing, the physical characteristic that the installed capacity of distributed photovoltaics remains unchanged in a short period of time is fully utilized to effectively suppress the jump in identification results caused by measurement noise and improve the temporal continuity of identification results.
[0033] 6. Closed-loop verification mechanism ensures the reliability of results: A reasonable capacity range based on transformer rated capacity, historical application data, solar resource level and penetration limit is set to automatically verify the identification results. When anomalies occur, data quality backtracking and re-identification are triggered, forming a closed-loop process of "identification-verification-correction" to ensure the engineering credibility of the output results.
[0034] 7. Edge-cloud collaborative architecture supports engineering applications: The lightweight model on the edge side meets real-time requirements, the high-precision model on the cloud ensures long-term accuracy, and the incremental synchronization mechanism ensures that the parameters on both sides are consistent. This architecture makes full use of the existing power distribution network communication infrastructure and can be deployed and implemented without large-scale transformation.
[0035] The technical effects achieved by this invention are as follows: 1. No dedicated measurement equipment required, reducing system upgrade costs: Existing distributed photovoltaic capacity identification methods mostly rely on the installation of dedicated measurement equipment such as smart meters and micro synchronous phasor measurement devices, which suffers from high investment, difficult construction, and low user cooperation. This invention utilizes only existing distribution network dispatch and control system measurement data, public meteorological service data, and power grid GIS topology data to achieve capacity identification through multi-source information fusion and intelligent algorithms. No hardware modifications to on-site equipment are required, significantly reducing the technology promotion threshold and engineering implementation costs.
[0036] 2. Multi-source data complementarity and fusion enhance identification accuracy and robustness: Single data sources have inherent limitations: measurement data is susceptible to bad data due to communication interruptions and sensor malfunctions; satellite cloud imagery has coarse spatial resolution but wide coverage; ground meteorological station data has high accuracy but is sparse; numerical weather prediction provides trend guidance but is limited in timeliness. This invention organically integrates three types of meteorological data with power grid measurement data and topology data through spatiotemporal alignment and standardization, fully leveraging the complementary advantages of each data source in terms of coverage, spatial accuracy, and temporal granularity. This effectively overcomes the interference of missing information or local anomalies on the identification results, significantly improving overall identification accuracy compared to single-data-source methods.
[0037] 3. Adaptive Partition Aggregation, Balancing Computational Efficiency and Identification Accuracy: High-penetration distribution networks have a large number of nodes. Identifying each node independently results in high computational complexity and neglects the electrical coupling relationships between nodes, easily leading to distorted results. This invention uses adaptive partitioning based on electrical distance and operational characteristic similarity, decomposing the large-scale network into several homogeneous sub-regions. While preserving the electrical characteristics and meteorological response patterns within each region, it significantly reduces the computational dimensionality through equivalent sequence aggregation. The number of sub-regions is automatically determined according to the power grid scale, eliminating the need for manual parameter tuning. This satisfies both online real-time requirements and ensures the engineering usability of the identification results.
[0038] 4. Meteorological-Load Joint Separation to Eliminate Non-PV Interference Factors: Besides PV output, net load also includes temperature-sensitive loads (such as air conditioning) and radiation-sensitive loads (such as lighting), which can easily lead to capacity identification errors if directly inverted. This invention establishes a multivariate correlation model between net load and meteorological factors through meteorological correlation analysis and partial correlation processing, effectively identifying and eliminating the aforementioned interference components. Simultaneously, a sliding time window is used to dynamically update model parameters, tracking the seasonal evolution of load characteristics, making PV output separation more accurate and the identification results closer to the actual installed capacity.
[0039] 5. Variable step size optimization and time constraints work together to ensure convergence stability and result continuity: Capacity identification is essentially a non-convex optimization problem, and traditional fixed step size gradient descent methods are prone to oscillations or getting trapped in local optima. This invention adaptively adjusts the step size based on the ratio of the residual change rate, performing a finer search when the descent rate slows down, and restoring the initial step size and adjusting the direction when oscillations occur, significantly improving convergence stability. Simultaneously, utilizing the physical characteristic that the installed capacity of distributed photovoltaic systems remains unchanged in the short term, a time consistency penalty term and exponential smoothing post-processing are introduced to effectively suppress jumps in identification values caused by measurement noise. The output results have good time continuity, facilitating tracking and analysis by dispatchers.
[0040] 6. Closed-loop verification mechanism to ensure the reliability of results: The identification algorithm is affected by multiple factors such as data quality, model assumptions, and parameter settings, which may lead to abnormal output. This invention sets a reasonable capacity range based on transformer capacity, application data, solar resource level, and upper limit of penetration rate to automatically verify the identification results; when an anomaly occurs, data quality backtracking is triggered, checking each link in measurement, meteorology, and zoning in sequence, and re-identifying after repair, forming a closed loop of "identification-verification-correction". This mechanism effectively isolates the impact of data faults on the final output, ensuring that the capacity information entering the scheduling decision system is true and reliable.
[0041] 7. Edge-Cloud Collaborative Architecture: Balancing Real-Time Response and Long-Term Accuracy: Distribution network dispatching demands both real-time performance and accuracy in capacity identification, but these two aspects conflict in terms of computing resources. This invention employs an edge-cloud collaborative architecture: a lightweight model is deployed at the edge, utilizing the local computing resources of the power supply station to achieve second-level real-time estimation, meeting the timeliness requirements of online monitoring; a high-precision model is deployed in the cloud, utilizing the computing power of the dispatch center for daily periodic verification and global optimization, ensuring long-term statistical accuracy. An incremental synchronization mechanism ensures parameter consistency on both sides. This architecture fully utilizes existing communication infrastructure and can be deployed and implemented without additional network investment.
[0042] The economic benefits that this invention can achieve are as follows: 1. Reduce distribution network operation and management costs: Accurate distributed photovoltaic capacity information is the foundation for distribution network dispatch planning, voltage and reactive power optimization, and fault location and isolation. The online identification results provided by this invention can be directly integrated into the existing dispatch control system, replacing manual on-site verification and ledger maintenance, reducing the frequency of business trips for maintenance personnel, and lowering the labor and time costs of distribution network operation and management.
[0043] 2. Enhance Distributed Energy Consumption Benefits: Based on real-time capacity identification results, dispatching departments can accurately assess the distributed photovoltaic (PV) absorption capacity of each region, scientifically guide the orderly grid connection of new PV installations, and avoid grid security issues caused by conservative power rationing or blind grid connection due to information gaps. This maximizes the space for distributed energy consumption while ensuring grid security, thereby improving the utilization rate of clean energy and power generation revenue.
[0044] 3. Reduce grid upgrade and transformation costs: Traditional distribution network planning often reserves redundant capacity based on maximum penetration rate, resulting in low investment efficiency. The online capacity monitoring data provided by this invention can be used to verify distribution network planning schemes, accurately identify photovoltaic hotspots and weak links, guide targeted grid upgrades and transformations, avoid over-investment, and improve the utilization efficiency of grid assets.
[0045] The social effects that this invention can achieve are as follows: 1. Supporting the Construction of New Power Systems: Distributed photovoltaic (PV) power is a crucial support for building new power systems and achieving "dual carbon" goals. This invention addresses the industry pain points of distributed PV being "invisible and unmanageable" in high-penetration scenarios, providing key state-sensing technologies for the transformation of distribution networks from "passive" to "active" forms, promoting coordinated interaction between power generation, grid, load, and storage, and driving the evolution of the power system towards a clean, low-carbon, safe, and efficient direction.
[0046] 2. Ensuring the safe and stable operation of the power grid: Large-scale and disorderly integration of distributed photovoltaic (PV) power can easily lead to safety issues such as voltage exceeding limits, backflow of power, and malfunctioning protection systems in the distribution network. This invention, through online capacity identification and anomaly early warning, enables dispatching departments to promptly grasp the actual scale of PV integration and take measures such as voltage regulation, power flow control, and orderly regulation in advance. This effectively prevents operational risks brought about by the high penetration rate of distributed energy and ensures the reliability of public power grid supply and the safety of user electricity consumption.
[0047] 3. Promote the standardized development of distributed photovoltaic power: Accurate capacity monitoring data can provide objective basis for the management of distributed photovoltaic power registration, subsidy verification, and tax collection, curb bad behaviors such as false reporting of capacity and illegal capacity expansion, promote the healthy and orderly development of the industry, and maintain a fair market competition environment. Attached Figure Description
[0048] Figure 1 This is a block diagram illustrating the principle of a distributed photovoltaic capacity online identification method based on multi-source data fusion according to the present invention. Figure 2 This is a structural block diagram of a distributed photovoltaic capacity online identification system based on multi-source data fusion, as described in this invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description is provided in conjunction with the accompanying drawings (reference 2). Figure 1 and Figure 2 The invention will be further described in detail below with reference to specific embodiments. Example 1: Multi-source data acquisition and preprocessing:
[0050] The study focuses on a 10kV distribution network area under the jurisdiction of a power supply company in a city in Henan Province. This area contains 86 load nodes, some of which have distributed photovoltaic access, but specific capacity information is missing.
[0051] 1. Multi-source data acquisition: The system accesses three types of data sources in parallel through the data access layer: Measurement data: The three-phase voltage amplitude, injected active power, and injected reactive power of each node are obtained from the Distribution Network Dispatch and Control System (DSCADA), with a sampling period of 15 minutes; Meteorological data: including irradiance data retrieved from meteorological satellite cloud images (spatial resolution 4km, temporal resolution 15 minutes), measured data from three automatic weather stations in the region (irradiance, temperature, cloud cover index, temporal resolution 10 minutes), and regional numerical weather forecast data (irradiance and temperature forecast for the next 24 hours, temporal resolution 1 hour). Topology data: Obtain node connection relationships, line impedance parameters, transformer tap information, and historical installation location information of distributed photovoltaic power from the power grid GIS system.
[0052] 2. Identification and repair of bad measurement data: An abnormal measurement value is identified by a residual detection method based on node power balance. The node power balance residual is calculated. If the residual exceeds three times the standard deviation of the historical residual of the node, the corresponding measurement value is determined to be an abnormal value.
[0053] For identified abnormal measurement values, the trend extrapolation value of historical data from the same period at the same node is used first for repair. Specifically, the measurement values at the same time 7 days before the node are taken, the average value is calculated, and the repair value is obtained by combining it with linear trend extrapolation.
[0054] If historical data for a node is missing (e.g., for a newly commissioned node), weighted interpolation of measurements from nodes in the same partition is used for repair. The weights of the weighted interpolation are determined based on the reciprocal of the electrical distance between nodes. The electrical distance is calculated by accumulating the line impedance parameters, with nodes that are closer together having a greater weight.
[0055] 3. Spatiotemporal alignment and standardization of meteorological data Meteorological data at different spatial resolutions are mapped to distribution network nodes using inverse distance weighted interpolation. For irradiance retrieval from satellite cloud images, the equivalent irradiance of the node is calculated by weighting the four pixel centers closest to the node using the inverse of their geographical distance. For measured data from ground meteorological stations, the data is also mapped to each node using the inverse of the geographical distance from the meteorological station.
[0056] Meteorological data at different time resolutions are aligned to the measurement data timestamp (15 minutes in total) using linear interpolation. For 10-minute data from ground meteorological stations, the values at the 15-minute time are obtained by linear interpolation of data from two adjacent 10-minute intervals. For 1-hour numerical weather prediction data, linear interpolation is also used to align the data to 15 minutes.
[0057] The aligned meteorological data are standardized to eliminate dimensional differences between different meteorological factors, which facilitates subsequent fusion analysis.
[0058] 4. Topology data node correlation verification: Two checks are performed: first, connectivity check, which uses a depth-first search algorithm to check topological connectivity and identify island nodes caused by incorrect switch status; second, electrical consistency check, which uses line impedance parameters and measurement data to calculate node voltage estimates and compares them with the measured voltages. If the deviation exceeds 5%, the line parameters are marked as potentially incorrect, triggering manual review.
[0059] After verification, a multi-source fusion dataset with a unified spatiotemporal benchmark is constructed. Each record contains a node number, timestamp, voltage measurement, power measurement, standardized multi-source meteorological data, and topological association identifier. Example 2: Extraction and Partition Aggregation of Running Status Features
[0060] 1. Feature extraction of net load power curve: For each node, the following net load power curve features are extracted based on the multi-source fusion dataset: Daily peak-valley difference rate: reflects the amplitude of daily load fluctuations; Average daily power: reflects the overall load level; Power fluctuation coefficient: reflects load stability; Power ramp rate: reflects the degree of drastic change in load.
[0061] 2. Voltage sensitivity feature extraction: Based on the line impedance parameters in the topology data, the voltage sensitivity of each node to active power injection and voltage sensitivity to reactive power injection are calculated to reflect the electrical position characteristics of the node in the power grid.
[0062] 3. Meteorological response feature extraction: Calculate the time-delay correlation coefficients between the net load power of each node and temperature, irradiance, and cloud cover index. By traversing the time delay range of 0 to 3 hours, determine the optimal time delay that maximizes the correlation coefficient and identify the meteorological sensitivity characteristics of the nodes.
[0063] 4. Adaptive partitioning The comprehensive similarity index between nodes is calculated. This index is composed of a weighted average of the reciprocal of the electrical distance and the cosine similarity of the feature vectors, with electrical distance accounting for 60% and feature similarity accounting for 40%. A clustering algorithm based on local density and relative distance is used for partitioning. Local density is determined by the number of neighboring nodes within a node's electrical distance threshold, which is the median of all node electrical distances. Relative distance is determined by the minimum electrical distance from a node to a node with higher local density. Nodes with significantly higher local density and relative distance than other nodes are selected as cluster centers, and the remaining nodes are assigned to the partition of the nearest higher-density node. In this embodiment, 86 nodes are automatically divided into 7 identification sub-regions, with each sub-region having between 8 and 16 nodes. Nodes within the same sub-region have similar electrical coupling characteristics and meteorological response patterns.
[0064] 5. Equivalent sequence aggregation: Within each sub-region, the equivalent net load series and equivalent meteorological series are aggregated and aligned by time. The equivalent net load series is the time-series sum of the net load of each node within the sub-region, and the equivalent meteorological series is the time-series average of the meteorological data of each node within the sub-region, weighted by load capacity. Example 3: Dynamic Decoupling of Distributed Photovoltaic Output:
[0065] 1. Establishment of the equivalent photovoltaic power output model: For each identified sub-region, an equivalent photovoltaic (PV) output model considering the coupling relationship of meteorological factors is established. The model uses equivalent irradiance, equivalent temperature, and cloud cover correction coefficients as inputs, and PV capacity as a linear scaling factor to describe the PV output characteristics under ideal conditions. The cloud cover correction coefficient is determined by looking up the cloud cover index in a table; the larger the cloud cover, the smaller the correction coefficient.
[0066] 2. Meteorological correlation analysis and output separation: A multivariate correlation model was established between the equivalent net load of a sub-region and meteorological factors. Partial correlation analysis was conducted to calculate the partial correlation coefficient between net load and irradiance under the condition of constant temperature, identifying irradiance-sensitive load components; and to calculate the partial correlation coefficient between net load and temperature under the condition of constant irradiance, identifying temperature-sensitive load components. The identified weather-sensitive load components were removed, and the remaining load components that decreased due to increased irradiance were separated into a photovoltaic (PV) output estimation sequence. PV output is represented as negative injected power in the net load, and the corresponding negative value was taken during separation.
[0067] 3. The sliding time window updates dynamically: A 30-day sliding time window is used to dynamically update model parameters. The window covers historical data for the most recent 30 days. After each identification period (1 day), the window slides forward 1 day to re-estimate the meteorological-sensitive load component and the base load component. This mechanism ensures that the model is always based on recent data, can track the seasonal and time-varying characteristics of load features, and adapt to changes in scenarios such as differences in air conditioning load between summer and winter. Example 4: Online Capacity Identification and Iterative Optimization
[0068] 1. Establishment of photovoltaic capacity-output mapping relationship: Based on the decoupled photovoltaic (PV) output estimation sequence and the corresponding meteorological irradiance data, a mapping relationship between PV capacity and output is established. Under the condition that the irradiance, temperature, and cloud cover correction coefficients are known, PV output and capacity exhibit a linear proportional relationship, with the proportionality coefficient determined by the meteorological conditions.
[0069] 2. Iterative solution using the variable step size gradient descent method: Using photovoltaic capacity as the optimization variable and minimizing the fitting residual between the photovoltaic power output estimation sequence and the standard photovoltaic power output curve as the objective function, the variable step size gradient descent method is used for iterative solution. During iteration, the step size is determined based on the ratio of the residual change rate of the current iteration step to the residual change rate of the previous iteration step. A preset decay threshold of 0.5 and a preset decay coefficient of 0.7 are set. When the above ratio is less than the preset decay threshold, it is determined that the residual descent rate has slowed down. The current step size is multiplied by the preset decay coefficient, and the step size is reduced to perform a fine search. When the residual changes in opposite directions in two consecutive iterations, it is determined that the search process has oscillated. The step size is restored to the initial value, and the search direction is adjusted to the opposite direction of the previous direction.
[0070] 3. Smooth correction of timing consistency constraints: A penalty term for capacity change between adjacent identification time periods is established in the objective function. This penalty term is proportional to the rate of capacity change. Incorporating this penalty term into the objective function creates a constrained optimization problem, suppressing drastic jumps in identification results between adjacent time periods. Meanwhile, exponential smoothing is used to post-process the identification results for multiple consecutive time periods (such as the last 7 days). The smoothing coefficient is set to 0.3, with the current identification result accounting for 30% and the historical smoothed result accounting for 70%. This fully utilizes the physical characteristic that the installed capacity of distributed photovoltaics remains unchanged in the short term and suppresses the jump in capacity identification value caused by measurement noise. After smoothing correction, the online identification results of the distributed photovoltaic capacity of each node are output. Example 5: Verification of Identification Results and Feedback on Anomalies
[0071] 1. Determining a reasonable capacity range: For each node, a reasonable capacity range is preset: the lower limit is 0, allowing nodes to operate without photovoltaic installations; the upper limit is determined comprehensively based on the following factors: the rated capacity of the transformer at the distribution network node (the maximum allowable active power after considering the power factor), historical installation data (if available, use a tolerance factor of 1.5), the capacity coefficient corresponding to the regional solar resource level, and the upper limit of distributed photovoltaic penetration rate (as stipulated in the distribution network dispatching regulations, such as 30%). The minimum value of the above calculated values is taken as the reasonable capacity upper limit.
[0072] 2. Identification result comparison and anomaly handling: The online identification results are compared with the preset reasonable capacity range: if the identification results are within the reasonable range, the verification is passed and the final result is directly output; if the identification results exceed the reasonable capacity range, the data quality backtracking verification mechanism is triggered: the measurement data quality is checked in sequence (bad data identification and repair are re-executed), the spatiotemporal alignment accuracy of meteorological data is checked (whether there are any abnormalities in the interpolation process), and the rationality of adaptive partitioning is checked (whether the boundary node assignment is appropriate). After the problems found are repaired, steps S1 to S4 are re-executed.
[0073] If the re-identification result still exceeds the reasonable range, or reaches the maximum number of iterations (set to 3), then mark the node as "identification abnormal", output the most recent identification result with an abnormality label, and prompt the scheduler to manually intervene and verify. Example 6: System Architecture and Engineering Applications
[0074] 1. System Composition: The multi-source data fusion distributed photovoltaic capacity online identification system in this embodiment is deployed in the dispatch and control center of the municipal power supply company and includes the following modules: Data access layer: Through the dispatch data network, it accesses the distribution network measurement system (IEC 60870-104 protocol), meteorological service system (RESTful API interface), and power grid GIS system (WebService interface) to achieve unified access and protocol conversion of multi-source data; Data preprocessing module: Deployed on the central server, it performs the multi-source data acquisition and preprocessing described in step S1, with a processing cycle of 15 minutes; Feature extraction and partitioning module: Deployed on the central server, it performs the running status feature extraction and partitioning aggregation described in step S2, and the partitioning results are updated once a day; Output decoupling module: Deployed on the central server, it performs the distributed photovoltaic output dynamic decoupling described in step S3, and the sliding window parameters are configurable; Capacity identification module: Adopts an edge-cloud collaborative architecture. The edge side is deployed at each power supply station, running a lightweight identification model to perform real-time capacity estimation with a response time of less than 5 seconds; the cloud side is deployed at the dispatch center, running a high-precision identification model, and performing regular verification and model updates every morning; the edge side and the cloud side maintain the consistency of identification parameters through an incremental synchronization mechanism. Result verification module: Deployed on the central server, it performs the identification result verification and anomaly feedback as described in step S5; Human-computer interaction interface: Based on WebGIS, it displays the online identification results, identification confidence level, anomaly markers and historical trend curves of the distributed photovoltaic capacity of each node.
[0075] 2. Engineering Applications: The online identification results of this system are integrated into the power distribution network dispatch and control system, and are specifically applied to: Distributed photovoltaic (PV) grid connection capacity assessment: Based on real-time capacity identification results, calculate the remaining grid connection capacity in each region to guide the orderly grid connection of new PV installations; Distribution network voltage over-limit early warning: Combining capacity identification results and output forecast, predict the risk of voltage rise during peak photovoltaic power generation periods, and adjust transformer taps or activate reactive power compensation in advance. Distributed energy dispatch planning: The identified capacity is used as a boundary condition input to the dispatch optimization model to improve the accuracy of photovoltaic output prediction and optimize unit combination and power flow distribution; Verification of distribution network planning schemes: Using long-term accumulated capacity identification data, we verify the photovoltaic penetration rate forecast for the planning year, identify weak links in the power grid, and guide the upgrading and transformation of the grid structure.
[0076] The above are merely preferred embodiments of the present invention, and the present invention is not limited to the specific embodiments described above. Those skilled in the art should understand that various modifications, equivalent substitutions, or improvements can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications, equivalent substitutions, or improvements should be covered within the protection scope of the present invention.
[0077] For example, in step S1, the sources of meteorological data can be expanded to include auxiliary data sources such as drone aerial cloud images and limited state information reported by distributed photovoltaic inverters; bad data repair methods can also use other technical means such as Kalman filtering and neural network interpolation to replace trend extrapolation and weighted interpolation.
[0078] In step S2, the adaptive partitioning algorithm can use other clustering methods such as K-means and spectral clustering to replace the clustering algorithm based on local density and relative distance; the feature vector can be constructed by increasing or decreasing the feature dimension according to the characteristics of the power distribution network.
[0079] In step S3, the length of the sliding time window can be adjusted to other values within the range of 15 to 60 days according to the differences in regional load characteristics; meteorological correlation analysis can be performed using other multivariate statistical methods such as principal component analysis and independent component analysis.
[0080] In step S4, the iterative optimization algorithm can use other optimization methods such as Newton's method, quasi-Newton's method, and genetic algorithm to replace the variable step size gradient descent method; the temporal consistency constraint can use other filtering methods such as moving average and Kalman smoothing to replace the exponential smoothing method.
[0081] In step S5, the factors for determining the reasonable capacity range can include constraints such as local subsidy policies and available roof area; the maximum number of iterations can be adjusted to other values within the range of 2 to 5 based on the system response requirements.
[0082] In terms of system architecture, the edge-cloud collaborative architecture can be adjusted to either a pure cloud centralized architecture or a pure edge distributed architecture, which can be flexibly selected according to the communication conditions and computing resources of the power distribution network.
Claims
1. A method for online identification of distributed photovoltaic capacity through multi-source data fusion, characterized in that, Includes the following steps: S1. Multi-source data acquisition and preprocessing: Acquire measurement data, meteorological data, and topology data of the target distribution network area; identify and repair bad data in the measurement data; perform spatiotemporal alignment and standardization processing on the meteorological data; verify node correlation in the topology data; and construct a multi-source fusion dataset with a unified spatiotemporal benchmark. The meteorological data includes irradiance data retrieved from satellite cloud images, measured data from ground meteorological stations, and numerical weather prediction data. The spatiotemporal alignment includes mapping meteorological data with different spatial resolutions to distribution network nodes using inverse distance weighting interpolation, and aligning meteorological data with different temporal resolutions to the measurement data timestamps using linear interpolation. S2. Operational status feature extraction and partition aggregation: Based on the multi-source fusion dataset, extract the net load power curve features, voltage sensitivity features and meteorological response features of each node, perform adaptive partitioning according to the electrical distance and feature similarity between nodes, divide the target distribution network area into several identification sub-regions, and aggregate in each sub-region to obtain the equivalent net load sequence and equivalent meteorological sequence; S3. Dynamic Decoupling of Distributed Photovoltaic Output: For each identified sub-region, an equivalent photovoltaic output model considering the coupling relationship of meteorological factors is established. Using the equivalent meteorological sequence and the equivalent net load sequence, the photovoltaic output is separated from the net load through meteorological correlation analysis to generate the photovoltaic output estimation sequence for each sub-region. The separation process uses a sliding time window to dynamically update the model parameters to adapt to the time-varying characteristics of the load. S4. Online Capacity Identification and Iterative Optimization: Based on the photovoltaic power output estimation sequence and the corresponding meteorological irradiance data, a photovoltaic capacity-output mapping relationship is established. A variable step-size gradient descent method is used for iterative solution. Photovoltaic capacity is the optimization variable, and minimizing the fitting residual between the photovoltaic power output estimation sequence and the standard photovoltaic power output curve is the objective function. The step size is determined based on the ratio of the residual change rate of the current iteration step to the residual change rate of the previous iteration step. When the ratio is less than a preset attenuation threshold, the current step size is multiplied by a preset attenuation coefficient. When the residual change directions are opposite in two consecutive iterations, the step size is restored to its initial value, the search direction is adjusted, and smoothing correction is performed through temporal consistency constraints of identification results in adjacent time periods. The online identification results of the distributed photovoltaic capacity of each node are then output. S5. Identification Result Verification and Anomaly Feedback: The online identification result is compared with the preset reasonable capacity range. If the identification result exceeds the reasonable capacity range, the data quality backtracking verification mechanism is triggered, and steps S1 to S4 are re-executed until the identification result meets the convergence condition or reaches the maximum number of iterations. The reasonable capacity range is determined comprehensively based on the rated capacity of the transformers at the distribution network nodes, historical application data, regional solar resource level, and the upper limit of distributed photovoltaic penetration rate.
2. The method for online identification of distributed photovoltaic capacity based on multi-source data fusion according to claim 1, characterized in that: In step S1, the identification and repair of bad data in the measurement data includes: using a residual detection method based on node power balance to identify abnormal measurement values; for the identified abnormal measurement values, priority is given to using the trend extrapolation value of the historical data of the same node for repair; if the historical data is missing, the weighted interpolation of the measurement values of the nodes in the same partition is used for repair; the weight of the weighted interpolation is determined according to the reciprocal of the electrical distance between nodes.
3. The method for online identification of distributed photovoltaic capacity based on multi-source data fusion according to claim 1, characterized in that: In step S2, the adaptive partitioning based on the electrical distance and feature similarity between nodes includes: calculating a comprehensive similarity index between nodes, which is composed of the inverse of the electrical distance and the cosine similarity of the feature vectors; using a clustering algorithm based on local density and relative distance for partitioning, wherein the local density is determined based on the number of neighboring nodes within the electrical distance threshold range of a node, and the relative distance is determined based on the minimum electrical distance from a node to a node with higher local density, so that nodes in the same sub-region have similar electrical coupling characteristics and meteorological response patterns, and the number of sub-regions is automatically determined according to the scale of the distribution network.
4. The method for online identification of distributed photovoltaic capacity based on multi-source data fusion according to claim 1, characterized in that: In step S4, the variable step size gradient descent method includes: using photovoltaic capacity as the optimization variable, minimizing the fitting residual between the photovoltaic power output estimation sequence and the standard photovoltaic power output curve as the objective function, the step size is determined according to the ratio of the residual change rate of the current iteration step to the residual change rate of the previous iteration step, when the ratio is less than a preset attenuation threshold, the current step size is multiplied by a preset attenuation coefficient; when the residual change directions are opposite in two consecutive iterations, the step size is restored to the initial value and the search direction is adjusted.
5. The method for online identification of distributed photovoltaic capacity based on multi-source data fusion according to claim 1, characterized in that: In step S4, the smoothing correction through the temporal consistency constraint of the identification results of adjacent time periods includes: establishing a capacity change penalty term for adjacent identification time periods, wherein the penalty term is proportional to the capacity change rate; adding the penalty term to the objective function to form an optimization problem with constraints; and using the exponential smoothing method to post-process the identification results of multiple consecutive time periods to suppress the jump in capacity identification values caused by measurement noise.
6. A distributed photovoltaic capacity online identification system based on multi-source data fusion, characterized in that... The system comprises a data access layer, a data preprocessing module, a feature extraction and partitioning module, an output decoupling module, a capacity identification module, and a result verification module. The data access layer is used to access multi-source data from the distribution network measurement system, meteorological service system, and power grid GIS system. The data preprocessing module is used to perform the multi-source data acquisition and preprocessing described in step S1 of claim 1. The feature extraction and partitioning module is used to perform the operating status feature extraction and partitioning aggregation described in step S2 of claim 1. The output decoupling module is used to perform the dynamic decoupling of distributed photovoltaic output described in step S3 of claim 1. The capacity identification module is used to perform the online capacity identification and iterative optimization described in step S4 of claim 1. The result verification module is used to perform the identification result verification and anomaly feedback described in step S5 of claim 1.
7. The distributed photovoltaic capacity online identification system based on multi-source data fusion according to claim 6, characterized in that: The capacity identification module adopts an edge-cloud collaborative architecture. A lightweight identification model is deployed on the edge to perform real-time capacity estimation, while a high-precision identification model is deployed on the cloud to perform periodic verification and model updates. The edge and cloud maintain the consistency of identification parameters through an incremental synchronization mechanism.
8. The method for online identification of distributed photovoltaic capacity based on multi-source data fusion according to claim 1, characterized in that: The method is applied to the distribution network dispatch and control system, and the online identification results are used for distributed photovoltaic acceptance capacity assessment, distribution network voltage over-limit early warning, distributed energy dispatch plan preparation, and distribution network planning scheme verification.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the distributed photovoltaic capacity online identification method based on multi-source data fusion as described in any one of claims 1 to 5.
10. An electronic device comprising a processor and a memory, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the distributed photovoltaic capacity online identification method based on multi-source data fusion as described in any one of claims 1 to 5.