A multi-source heterogeneous distributed photovoltaic data processing method and device and electronic equipment
By generating a tree-structured ledger, linking meteorological data, and performing rolling corrections and multiple rounds of abnormal data screening, the problem of large computational load and poor results in completing distributed photovoltaic data resources has been solved, achieving efficient data quality improvement and full integration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID BEIJING ELECTRIC POWER CO
- Filing Date
- 2023-11-17
- Publication Date
- 2026-07-21
AI Technical Summary
Existing distributed photovoltaic data resource completion schemes involve large computational loads and poor completion results, affecting data quality and the normal operation of application functions.
By acquiring the topological hierarchy of distributed photovoltaic (PV) systems to distribution transformers, feeders, busbars, substations, and regions, a tree-structured ledger is generated. Based on geographical location, power measurement data and meteorological data are correlated, and rolling correction and multi-source meteorological data fusion are performed. Distributed PV operation rules and outlier detection algorithms are used to identify abnormal data, and multi-round screening and completion correction strategies are adopted, including Lagrange interpolation and multi-site weighted methods to complete missing outliers.
It improved the accuracy of meteorological data in numerical weather forecasting, refined the criteria for screening abnormal data, enhanced the quality of distributed photovoltaic-related resource data, avoided the adverse effects of missing abnormal data on application functions, and achieved full integration and efficient completion of distributed photovoltaic data.
Smart Images

Figure CN117555887B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy technology, and in particular to a multi-source heterogeneous distributed photovoltaic data processing method, apparatus, and electronic equipment. Background Technology
[0002] In recent years, the proportion of new energy sources in terms of power and capacity has been increasing, gradually taking a mainstream position. With the increasing maturity of distributed generation and its system integration technology, distributed generation, with its advantages of low pollution, high energy efficiency, flexible installation locations, short construction periods, and simple maintenance, is being used more and more widely, gradually becoming another effective way to develop and utilize renewable energy. Compared with centralized power generation that relies on long-distance transmission and distribution, distributed generation uses local energy and consumes it locally, reducing line losses and transmission and transformation investment and operating costs associated with large-scale long-distance power transmission. It also improves the power quality at the end of the distribution network and reduces line losses. Distributed generation serves as a backup for the main power grid, enhancing the reliability and flexibility of power supply.
[0003] The power distribution network system is the transmission endpoint of the entire power system, characterized by numerous points of origin and wide distribution, and susceptibility to environmental influences. The small capacity and large number of distributed power sources significantly increase the complexity of system scheduling. With the integration of a large number of distributed power sources, the requirements for the integration, processing, and fusion of multi-source heterogeneous data from the distribution network are becoming increasingly stringent. Furthermore, with the gradual improvement of supporting technologies such as sensing and measurement, and information communication in new active distribution networks, a large amount of multi-source heterogeneous data is generated. Affected by issues such as the data acquisition environment, equipment failures, and communication defects, the large amount of distributed resource data collected may contain certain degrees of error and omissions, affecting data quality and consequently impacting the normal operation of various data-driven applications.
[0004] Existing distributed photovoltaic data resource integration technologies mainly include two aspects: First, the identification of abnormal data at the source end. This aspect currently primarily includes judgment methods based on data distribution indicators, such as the standard deviation method and the box plot method. The standard deviation method assumes that all samples in the sample set are normally distributed, calculates the standard deviation of the sample set, and identifies samples whose mean falls outside three times the standard deviation as outliers. The box plot method sets the upper quartile, lower quartile, and threshold coefficients for the sample set; values exceeding the threshold are considered outliers. Anomaly identification also includes density-based clustering methods, such as DBSCAN, and unsupervised learning methods based on decision trees, such as isolated forests and random forests. Density clustering outlier detection methods typically set a distance threshold and use density clustering to find isolated points that do not belong to any cluster, thus identifying outliers. Decision tree-based outlier detection methods directly and explicitly isolate outliers using decision trees. Existing distributed photovoltaic (PV) data resource completion technologies primarily address missing outliers in the source data. These methods mainly fall into two categories: first, using representative attributes of the distributed PV resource dataset, such as mode, mean, and typical values, to replace missing outliers or to substitute them based on neighboring data; second, fitting missing outliers using first-order or second-order interpolation, regression models, or decision tree models, thereby completing the missing outlier data at the distributed resource data source. However, existing distributed PV data resource completion schemes are computationally intensive and produce poor data completion results. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-source heterogeneous distributed photovoltaic data processing method, apparatus, and electronic device that can solve the problems of large computational load and poor data completion effect in existing distributed photovoltaic data resource completion schemes.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for processing multi-source heterogeneous distributed photovoltaic data, including: Obtain the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regions; Based on the aforementioned topological hierarchy and the ledger data of distributed photovoltaic power generation, a tree-shaped ledger of distributed photovoltaic power generation is generated. Based on the tree-shaped ledger and the power measurement data of distributed photovoltaics, the power measurement data of distributed photovoltaics at each level are determined; Based on geographical location, the distributed photovoltaic power measurement data at each level are correlated with meteorological data; Rolling correction is performed on the gridded meteorological data in the meteorological data to enable the fusion of multi-source meteorological data; Based on the distributed photovoltaic operation rules, the distributed photovoltaic power measurement data of each level are identified for the first time to obtain normal distributed photovoltaic power measurement data. The normal distributed photovoltaic power measurement data is identified a second time based on the outlier detection algorithm to obtain abnormal distributed photovoltaic power measurement data. The abnormal data in the distributed photovoltaic power measurement were supplemented and corrected.
[0007] Optionally, the step of associating the distributed photovoltaic power measurement data at each level with meteorological data based on geographical location includes: Load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; Based on the size of the square grid point in the numerical weather forecast and the coordinates of any corner of the grid point, determine the coordinates of the four corners of the grid point; Determine whether distributed photovoltaic systems are within the grid area; The distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid area is calculated, and the nearest meteorological monitoring station is used as the predictive meteorological data source for the distributed PV equipment, so as to realize the correlation between distributed PV power measurement data and meteorological data.
[0008] Optionally, the step of performing rolling correction on the gridded meteorological data in the meteorological data to enable multi-source meteorological data fusion includes: The time interval data monitored by meteorological monitoring devices are converted into data with the same time interval as numerical weather forecasts. The weighted average error value is calculated based on the real-time time-limited data monitored by the meteorological monitoring device on that day. The grid meteorological data in the meteorological data is offset based on the weighted average error value to achieve correction.
[0009] Optionally, the step of supplementing and correcting the abnormal data in the distributed photovoltaic power measurement includes: Determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; Based on the preset completion and correction strategy corresponding to the target time period, the abnormal data of the distributed photovoltaic power measurement are completed and corrected.
[0010] Optionally, the step of completing and correcting the abnormal distributed photovoltaic power measurement data according to the preset completion and correction strategy corresponding to the target time period includes: If the target time period is a specified nighttime period, the missing abnormal point data in the abnormal data of the distributed photovoltaic power measurement will be set to 0; When the target time period is the first time period, sample distributed photovoltaics that meet the conditions are screened based on the correlation of the daily power curves of distributed photovoltaics in the same region. Based on the number of sample distributed photovoltaic systems that meet the conditions, estimate the distributed photovoltaic power generation power of the distributed photovoltaic power measurement anomaly data that is missing abnormal data. In the absence of a qualified sample distributed photovoltaic system, the Lagrange interpolation method is used to complete and correct the abnormal power measurement data of the distributed photovoltaic system.
[0011] This invention also provides a multi-source heterogeneous distributed photovoltaic data processing device, comprising: The acquisition module is used to acquire the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regions. The generation module is used to generate a tree-shaped ledger of distributed photovoltaics based on the topological hierarchy and the ledger data of distributed photovoltaics. The determination module is used to determine the distributed photovoltaic power measurement data at each level based on the tree-shaped ledger and the power measurement data of the distributed photovoltaic system. The association module is used to associate the distributed photovoltaic power measurement data of each level with meteorological data based on geographical location; The correction module is used to perform rolling correction on the grid meteorological data in the meteorological data so as to fuse the multi-source meteorological data; The first identification module is used to perform the first identification of the distributed photovoltaic power measurement data of each level based on the distributed photovoltaic operation rules, so as to obtain normal distributed photovoltaic power measurement data. The second identification module is used to perform a second identification on the normal distributed photovoltaic power measurement data based on the outlier detection algorithm to obtain abnormal distributed photovoltaic power measurement data. The correction module is used to complete and correct the abnormal data of the distributed photovoltaic power measurement.
[0012] Optionally, the association module includes: The first submodule is used to load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; The second submodule is used to determine the coordinates of the four corners of the grid point based on the size of the square grid point in the numerical weather forecast and the coordinates of any corner of the grid point; The third submodule is used to determine whether the distributed photovoltaic system is within the grid range; The fourth submodule is used to calculate the distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid range, and to select the nearest meteorological monitoring station as the predictive meteorological data source for the distributed PV system, so as to realize the correlation between distributed PV power measurement data and meteorological data.
[0013] Optionally, the correction module includes: The fifth submodule is used to convert the time interval data monitored by the meteorological monitoring device into data with the same time interval as the numerical weather forecast. The sixth submodule is used to calculate the weighted average error value based on the real-time time-limited data monitored by the meteorological monitoring device on that day. The seventh submodule is used to offset the grid meteorological data in the meteorological data based on the weighted average error value in order to achieve correction.
[0014] Optionally, the correction module includes: The eighth submodule is used to determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; The ninth submodule is used to complete and correct the abnormal data of the distributed photovoltaic power measurement according to the preset completion and correction strategy corresponding to the target time period.
[0015] Optionally, the ninth submodule is specifically used for: If the target time period is a specified nighttime period, the missing abnormal point data in the abnormal data of the distributed photovoltaic power measurement will be set to 0; When the target time period is the first time period, sample distributed photovoltaics that meet the conditions are screened based on the correlation of the daily power curves of distributed photovoltaics in the same region. Based on the number of sample distributed photovoltaic systems that meet the conditions, estimate the distributed photovoltaic power generation power of the distributed photovoltaic power measurement anomaly data that is missing abnormal data. In the absence of a qualified sample distributed photovoltaic system, the Lagrange interpolation method is used to complete and correct the abnormal power measurement data of the distributed photovoltaic system.
[0016] This invention provides an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of any of the above-described multi-source heterogeneous distributed photovoltaic data processing methods.
[0017] This invention provides a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of any of the above-described multi-source heterogeneous distributed photovoltaic data processing methods.
[0018] This invention provides a multi-source heterogeneous distributed photovoltaic (PV) data processing scheme, which acquires the topological hierarchy of distributed PV systems to distribution transformers, feeders, busbars, substations, and regions; generates a tree-structured ledger of distributed PV systems based on the topological hierarchy and the ledger data of distributed PV systems; determines the power measurement data of distributed PV systems at each level based on the tree-structured ledger and the power measurement data of distributed PV systems; associates the power measurement data of distributed PV systems at each level with meteorological data based on geographical location; performs rolling correction on the grid meteorological data in the meteorological data to achieve multi-source meteorological data fusion; performs a first identification of the power measurement data of distributed PV systems at each level based on the distributed PV operating rules to obtain normal distributed PV power measurement data; performs a second identification of the normal distributed PV power measurement data based on an outlier detection algorithm to obtain abnormal distributed PV power measurement data; and completes and corrects the abnormal distributed PV power measurement data. The multi-source heterogeneous distributed photovoltaic (PV) data processing scheme provided in this invention has two aspects. First, it integrates multi-source heterogeneous data such as numerical weather forecasts, meteorological data from meteorological monitoring equipment, and distributed PV power data to improve the accuracy of gridded meteorological data from numerical weather forecasts and achieve full integration of distributed PV-related resource data. Second, based on distributed PV operation rules and outlier detection algorithms, it identifies and filters outliers in distributed PV power data in two rounds. The outlier data screening criteria are more comprehensive, avoiding the adverse effects of missing outlier data on distributed PV-related application functions, thereby improving the data quality of the supplemented and corrected distributed PV-related resource data. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the steps of a multi-source heterogeneous distributed photovoltaic data processing method according to an embodiment of this application; Figure 2 This is a schematic diagram of a tree-shaped ledger; Figure 3 This is a flowchart illustrating the process of handling outliers in distributed photovoltaic systems. Figure 4 This is a schematic diagram illustrating the principle of the multi-site weighted method; Figure 5 This is a schematic diagram of the process for completing and correcting missing and abnormal data in distributed photovoltaic systems. Figure 6 This is a structural block diagram illustrating an embodiment of the multi-source heterogeneous distributed photovoltaic data processing device of this application; Figure 7 This is a structural block diagram illustrating an embodiment of an electronic device according to this application. Detailed Implementation
[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0021] This application discloses a method for integrating, fusing, and supplementing multi-source heterogeneous power distribution and distributed photovoltaic (PV) data resources. This method integrates and fuses distributed resource data such as measurement, topology, and meteorological data from distributed PV connected to the distribution network. Considering the basic operating principles and rules of distributed PV, it employs the Local Outlier Factor (LOF) algorithm and a two-round outlier screening method to detect daily power anomalies in distributed PV data. Considering the daily power operation characteristics of distributed PV and the correlation of daily power in different regions, it supplements missing outlier data. Based on surrounding sample distributed PV data, it uses a multi-site weighted method or a capacity reduction algorithm to supplement missing outliers. If no sample distributed PV is available in the vicinity, Lagrange interpolation is used for supplementation. This method for integrating, fusing, and supplementing multi-source heterogeneous power distribution and distributed PV data resources for large-scale distributed PV integration into the distribution network achieves the integration and fusion of multi-source heterogeneous distributed PV resource data, identifies missing outliers in distributed PV, and supplements and corrects these outliers, thereby improving the data quality of distributed PV resource data.
[0022] The multi-source heterogeneous distributed photovoltaic data processing scheme provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios. (See attached figures.) Figure 1 As shown, the multi-source heterogeneous distributed photovoltaic data processing method of this application embodiment includes the following steps: Step 101: Obtain the topological hierarchy of distributed photovoltaic power generation to distribution transformers, feeders, busbars, substations, and regions.
[0023] First, the topological hierarchy from distributed photovoltaic power generation to distribution transformers, feeders, busbars, substations, and regional levels is obtained. The original data structure of the topological hierarchy is shown in Table 1 below:
[0024] As shown in the table, the above data structure is used to describe the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regional levels, and is divided into three cases: For distribution transformers that are not connected to distributed photovoltaic systems, the distributed photovoltaic column should be written as "Null". One distribution transformer corresponds to one record, such as serial number 1 in the example. For medium-voltage distributed photovoltaic systems with direct T-type feeders, write "Null" in the distribution transformer column. One photovoltaic system corresponds to one record, such as serial number 2 in the example. For low-voltage distributed photovoltaic systems connected to the distribution transformer area, the system maintains the identification of the distribution transformer, feeder, busbar, substation, and region. One photovoltaic system corresponds to one record, such as serial number 3 in the example.
[0025] Step 102: Generate a tree-shaped ledger of distributed photovoltaics based on the topological hierarchy and the ledger data of distributed photovoltaics.
[0026] In practical implementation, a tree diagram of the relationship between devices at each level within the region can be formed based on the data structure of the topological hierarchy. At the same time, based on the ledger data such as the type and capacity of distributed photovoltaics, a tree ledger of the region, substation, busbar, feeder, distribution transformer, and distributed photovoltaics can be integrated.
[0027] Step 103: Based on the tree-structured ledger and the power measurement data of distributed photovoltaics, determine the power measurement data of distributed photovoltaics at each level.
[0028] Step 104: Based on geographical location, correlate the distributed photovoltaic power measurement data at each level with meteorological data.
[0029] After obtaining the distributed photovoltaic power measurement data at each level, it is necessary to understand the relationship between numerical weather forecast data, meteorological monitoring data, and distributed photovoltaic data to achieve the fusion of distributed photovoltaic ledgers, operation data, and heterogeneous meteorological data. The specific method is to consider the geographical locations of each grid point of the numerical weather forecast, meteorological monitoring stations, and distributed photovoltaic data, and associate meteorological data with distributed photovoltaic data based on the principle of the shortest Euclidean distance between geographical coordinates.
[0030] An optional method for correlating the distributed photovoltaic power measurement data at each level with meteorological data based on geographical location is as follows: First, load the latitude and longitude coordinates of the grid points, meteorological monitoring stations, and distributed photovoltaic systems; Secondly, based on the size of the square grid points in the numerical weather forecast and the coordinates of any corner of the grid point, the coordinates of the four corners of the grid point are determined; Next, determine whether the distributed photovoltaic system is within the grid area; Finally, the distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid area is calculated, and the nearest meteorological monitoring station is selected as the predictive meteorological data source for the distributed PV equipment, so as to realize the correlation between distributed PV power measurement data and meteorological data.
[0031] Step 105: Perform rolling correction on the gridded meteorological data in the meteorological data to enable the fusion of multi-source meteorological data.
[0032] Multi-source meteorological data fusion mainly considers the fusion and correction of meteorological data from numerical weather prediction grids and real-time meteorological data collected by meteorological devices. Since numerical weather prediction data is updated daily for 3.5 days with 15-minute intervals, and meteorological data monitored by meteorological devices is updated with 5-minute intervals, the grid meteorological data can be rolled and corrected based on the data from the meteorological monitoring devices, which are updated every 5 minutes, thereby achieving the fusion of multi-source meteorological data.
[0033] An optional method for performing rolling correction on gridded meteorological data in meteorological data to facilitate the fusion of multi-source meteorological data is as follows: The time interval data monitored by the meteorological monitoring device is converted into data with the same time interval as the numerical weather forecast; the weighted average error value is calculated based on the real-time period data monitored by the meteorological monitoring device on the same day; the grid meteorological data in the meteorological data is offset based on the weighted average error value to achieve correction.
[0034] Step 106: Based on the distributed photovoltaic operation rules, the distributed photovoltaic power measurement data at each level are identified for the first time to obtain normal distributed photovoltaic power measurement data.
[0035] In this embodiment of the application, in order to ensure the data quality of distributed photovoltaic power data, the power measurement data needs to be cleaned. First, it is necessary to identify abnormal data. The identification of abnormal data includes two identifications: abnormal data identification based on the distributed photovoltaic operation rules (i.e., the first identification) and abnormal data identification based on the outlier detection algorithm (i.e., the second identification).
[0036] The basic principles and operating rules of distributed photovoltaic systems related to anomaly data identification are as follows: a. Photovoltaic power generation directly converts light energy into electrical energy using the photovoltaic effect at the semiconductor interface; the source of the light energy is the sun. Therefore, some abnormal data can be identified based on the basic knowledge of sunrise and sunset and specific weather conditions. b. The capacity of distributed photovoltaic power generation is its maximum power output; c. During the transmission of distributed photovoltaic measurement data, issues such as packet loss during communication can cause the measurement data to remain unchanged for extended periods. d. Distributed photovoltaic systems are power generation devices that only generate electricity and do not have an electrical load.
[0037] Therefore, the criteria for filtering abnormal data from distributed photovoltaic systems based on operational rules are shown in Table 2 below:
[0038] Step 107: Based on the outlier detection algorithm, the normal distributed photovoltaic power measurement data is identified a second time to obtain abnormal distributed photovoltaic power measurement data.
[0039] After identification based on distributed photovoltaic operation rules, the remaining normal data undergoes a second identification using a density-based outlier identification algorithm—the Local Outlier Factor (LOF) algorithm—to obtain the final outlier data identification results. The core idea of the LOF algorithm is to compare the density around an object with the density of its neighborhood. The higher the local outlier factor of an object, the more it deviates from its surrounding points. Once a certain threshold is reached, it is identified as an outlier.
[0040] Step 108: Complete and correct any abnormal data in the distributed photovoltaic power measurement.
[0041] An optional method for completing and correcting anomalous data in distributed photovoltaic power measurement is as follows: Determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; and complete and correct the abnormal data in the distributed photovoltaic power measurement according to the preset completion and correction strategy corresponding to the target time period.
[0042] More specifically, based on the preset completion and correction strategy corresponding to the target time period, the method for completing and correcting abnormal data in distributed photovoltaic power measurement can be as follows: When the target time period is a specified nighttime period, the missing outlier data in the abnormal data of distributed photovoltaic power measurement will be set to 0; the specified nighttime period can be set to 0:00-4:00 and 21:00-23:45. With the target time period being the first time period, sample distributed photovoltaics in the surrounding area that meet the conditions are screened based on the correlation of the daily power curves of distributed photovoltaics in the same region; the first time period can be set to 4:00-21:00. Based on the number of qualified sample distributed photovoltaic (PV) systems, the power generation of distributed PV systems missing data in the abnormal power measurement data is estimated. In the absence of qualified sample distributed PV systems, the Lagrange interpolation method is used to complete and correct the abnormal power measurement data of distributed PV systems.
[0043] The multi-source heterogeneous distributed photovoltaic data processing provided in this application realizes the integration and fusion of distributed resource data such as measurement, topology, numerical weather prediction grid, and meteorological data from meteorological monitoring devices for distributed photovoltaic access to the distribution network. It uses the measured data from meteorological monitoring devices to continuously correct the grid meteorological data of numerical weather prediction, and adopts a two-round screening method of distributed photovoltaic operation rules + outlier detection algorithm LOF to detect abnormal data. It makes full use of the correlation of regional distributed photovoltaics and supplements the missing abnormal data of distributed photovoltaics based on the surrounding sample distributed photovoltaic data.
[0044] The multi-source heterogeneous distributed photovoltaic (PV) data processing scheme provided in this application has three aspects: First, it fuses multi-source heterogeneous data such as numerical weather forecasts, meteorological data from meteorological monitoring equipment, and distributed PV power data to improve the accuracy of gridded meteorological data from numerical weather forecasts and achieve full integration of distributed PV-related resource data. Second, based on distributed PV operation rules and outlier detection algorithms, it identifies and filters outliers in distributed PV power data in two rounds, resulting in more comprehensive outlier screening criteria and avoiding adverse effects of missing outlier data on distributed PV-related application functions. Third, it addresses the issue of... By fully utilizing the daily power correlation among distributed photovoltaic (PV) systems in the same region, and using PV power data from surrounding model stations, the missing outliers in distributed PV systems can be supplemented using a multi-site weighting method or a capacity discounting method, thereby improving the data quality of distributed PV-related resource data.
[0045] The following example illustrates the multi-source heterogeneous distributed photovoltaic data processing method provided in this application.
[0046] The multi-source heterogeneous distributed photovoltaic data processing method provided in this specific example includes the following steps: Step 1: Data integration and fusion of multi-source heterogeneous distributed photovoltaic resources.
[0047] This project aims to integrate and fuse multi-source heterogeneous data, including resource, asset, measurement, topology, and graphical data, as well as distributed resource data such as meteorological data, within the distribution network connected to distributed photovoltaic (PV) grids. A tree-structured ledger is established based on the topological hierarchy of distributed PV-distribution transformer-feeder-busbar-substation-region. Based on this tree-structured ledger and distributed PV power measurements, power measurements at each level are generated. Regarding meteorological data, the project integrates multi-source meteorological data, including meteorological data collected by meteorological monitoring devices in pilot areas and gridded meteorological data generated by numerical weather prediction. Based on the geographical location of equipment, meteorological monitoring devices, and grids, the project matches gridded meteorological data from numerical weather prediction with the equipment, achieving the fusion of heterogeneous data from distributed PV operation data, ledger data, and meteorological data. Specifically, this includes the following three points: (1) Establish a hierarchical tree-structured ledger to integrate distributed photovoltaic power measurement First, the topological hierarchy from distributed photovoltaic (PV) to distribution transformers, feeders, buses, substations, and regional levels is obtained. The original data structure of the topological hierarchy is shown in Table 1. As shown in Table 1, the above data structure is used to describe the topological hierarchy from distributed PV to distribution transformers, feeders, buses, substations, and regional levels, and is divided into three cases: For distribution transformers that are not connected to distributed photovoltaic systems, the distributed photovoltaic column should be written as "Null". One distribution transformer corresponds to one record, such as serial number 1 in the example. For medium-voltage distributed photovoltaic systems with direct T-type feeders, write "Null" in the distribution transformer column. One photovoltaic system corresponds to one record, such as serial number 2 in the example. For low-voltage distributed photovoltaic systems connected to the distribution transformer area, the system maintains the identification of the distribution transformer, feeder, busbar, substation, and region. One photovoltaic system corresponds to one record, such as serial number 3 in the example.
[0048] Based on a topological hierarchical data structure, a tree-like relationship diagram of devices at each level within the region is formed. Simultaneously, based on ledger data such as the type and capacity of distributed photovoltaic systems, a tree-like ledger is integrated for the region, substation, busbar, feeder, distribution transformer, and distributed photovoltaic systems. A schematic diagram of the tree-like ledger is shown below. Figure 2 As shown: Finally, based on the data ledger and the power measurement data of distributed photovoltaics, the measurement data of all equipment at each level were obtained.
[0049] (2) Integration of heterogeneous data on distributed photovoltaic ledgers, operation, and meteorology After obtaining the distributed photovoltaic (PV) power measurement data at each level, it is necessary to understand the relationship between numerical weather prediction data, meteorological monitoring data, and distributed PV, and to achieve the fusion of heterogeneous data on distributed PV records, operation, and meteorology. The specific method involves considering the geographical locations of each grid point in the numerical weather prediction, meteorological monitoring stations, and distributed PV, and using the principle of shortest Euclidean distance in geographical coordinates to associate the meteorological data with the distributed PV. The specific steps are as follows: ① Load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; ② Based on the size of the square grid point in the numerical weather forecast and the coordinates of a corner of the grid point, such as 10km×10km, the accurate coordinates of the four corners of the grid point are obtained; ③ Determine whether the distributed photovoltaic system is within the grid area; that is, the geographical coordinates of the distributed photovoltaic system must simultaneously satisfy both conditions a and b:
[0050] In the formula, E represents the longitude of the distributed photovoltaic system. and Indicates the longitude of the east and west boundaries of the grid points. and This indicates the latitude of the north and south boundaries of the grid point.
[0051] ④ Calculate the distance between each distributed photovoltaic system and each meteorological monitoring station, sort them in ascending order of distance, and take the nearest meteorological monitoring station as the predictive meteorological data source for this photovoltaic system.
[0052] (3) Fusion of multi-source meteorological data Multi-source meteorological data fusion primarily considers the fusion and correction of meteorological data from numerical weather prediction grids and real-time meteorological data collected by meteorological devices. Since numerical weather prediction data is updated daily for 3.5 days at 15-minute intervals, while meteorological data monitored by meteorological devices is updated at 5-minute intervals, within the grid containing the meteorological monitoring device, rolling correction can be performed on the grid meteorological data based on the 5-minute updated data from the monitoring device, thereby achieving multi-source meteorological data fusion. The specific correction method is as follows: ① Convert the 5-minute time interval data from the meteorological monitoring device into data with a 15-minute interval, similar to numerical weather prediction. The specific method is as follows:
[0053] In the formula, This represents the meteorological monitoring data at time t after transformation, with intervals of 15 minutes. , , This represents the meteorological data at 5-minute intervals before the transformation at times t-1, t, and t+1. w represents the weight, with the weight at time t being greater than that of the two points before and after it.
[0054] ② Based on meteorological monitoring data, the grid meteorological data of numerical weather forecasts is corrected in real time every hour. Specifically, the weighted average error value is calculated based on the real-time meteorological data from the meteorological monitoring device for the day. Subsequent numerical weather forecast meteorological data are then offset according to the average error value to achieve correction. The offset calculation method is as follows:
[0055] In the formula, Represents the weight vector. The sum of all elements in the array is 1. For meteorological monitoring data, This is numerical weather forecast grid data, where t is the current time.
[0056] This step involves matching the relationships between numerical weather prediction grid points, meteorological monitoring equipment, and distributed photovoltaic (PV) devices based on geographical location, thereby achieving heterogeneous data fusion of meteorological data and PV power generation data. It also uses measured data from meteorological monitoring devices to continuously correct the predicted meteorological data from the numerical weather prediction grid, achieving multi-source meteorological data fusion. This improves the accuracy of predicted meteorological data from numerical weather prediction grid points and enables full integration of distributed PV-related resource data.
[0057] Step 2: Identification of abnormal data in distributed photovoltaic systems.
[0058] To ensure the data quality of distributed photovoltaic power data, data cleaning is required for power measurement. The first step is to identify abnormal data, which involves two steps: (1) Abnormal data identification based on distributed photovoltaic operation rules The basic principles and operating rules of distributed photovoltaic systems related to anomaly data identification are as follows: a. Photovoltaic power generation directly converts light energy into electrical energy using the photovoltaic effect at the semiconductor interface; the source of the light energy is the sun. Therefore, some abnormal data can be identified based on the basic knowledge of sunrise and sunset and specific weather conditions. b. The capacity of distributed photovoltaic power generation is its maximum power output; c. During the transmission of distributed photovoltaic measurement data, issues such as packet loss during communication can cause the measurement data to remain unchanged for extended periods. d. Distributed photovoltaic systems are power generation devices that only generate electricity and do not have an electrical load.
[0059] Therefore, the criteria for filtering abnormal distributed photovoltaic data based on operating rules are shown in Table 2.
[0060] (2) Anomaly data identification based on outlier detection algorithm After identification based on distributed photovoltaic operation rules, the remaining normal data undergoes a second round of identification using a density-based outlier identification algorithm—the Local Outlier Factor (LOF) algorithm—to obtain the final outlier data identification results. The core idea of the LOF algorithm is to compare the density around an object with the density of its neighborhood. The higher the local outlier factor of an object, the more it deviates from its surrounding points, and when it reaches a certain threshold, it is identified as an outlier.
[0061] A flowchart illustrating the process of handling outliers in distributed photovoltaic systems is shown below. Figure 3 As shown, it includes the following nine steps ①-⑨: ① Load the data for 96 points of the distributed photovoltaic system for the day; ② Based on the distributed photovoltaic operation rules, the first round of abnormal data screening was carried out to obtain abnormal data point part1; ③ Load the distributed photovoltaic power generation data of the previous 30 days and combine it with the normal data points after the first round of screening to perform LOF outlier detection. The 30-day distributed photovoltaic power generation data is the basic data source for detection. ④ Define the LOF algorithm parameters k and the LOF threshold; ⑤ Calculate the k-nearest neighbor distance for each data point P. and k-distance neighborhood The k-nearest neighbor distance is the distance between sample P and its k-th nearest neighbor. The k-distance neighborhood is a circle centered at P. The neighborhood is defined by a radius, and the distance is calculated using Euclidean distance.
[0062] ⑥ Calculate the k-local reachability density for each data point P. The calculation method is as follows:
[0063] In the formula, Let P be the reachable distance from sample O, and take... The larger of the Euclidean distances between P and O.
[0064] ⑦ Calculate the k-local outlier factor for each data point P. The calculation method is as follows:
[0065] ⑧ Compare the k-local outliers for each data point P. ,like If so, it is an outlier; ⑨ The union of the abnormal data part1 based on the distributed photovoltaic operation rules and the abnormal data part2 obtained based on the LOF algorithm is the current abnormal data detection result of the distributed photovoltaic system.
[0066] This step considers the basic operating principles and rules of distributed photovoltaic (PV) systems. It employs a two-round outlier screening method based on PV operating rules and the Local Outlier Factor (LOF) algorithm to detect daily power anomalies in PV systems. The outlier points in the PV power data are identified and screened in two rounds, resulting in a more comprehensive outlier screening standard and avoiding the adverse effects of missing outlier data on PV-related application functions.
[0067] Step 3: Complete missing and abnormal data for distributed photovoltaic systems.
[0068] After identifying the missing anomalies in the distributed photovoltaic (PV) data for the day, it is necessary to complete and correct these anomalies. A schematic diagram of the process for completing and correcting missing anomaly data in distributed PV is shown below. Figure 5As shown, the overall completion and modification is carried out in three major steps, and each major step is further divided into at least one minor step: (1) Complete and correct the night-time distributed photovoltaic power generation values. For the distributed photovoltaic data from 0:00 to 4:00 and from 21:00 to 23:45 with missing or abnormal points, directly set them to 0.
[0069] (2) For the distributed photovoltaic power generation data from 4:00 to 21:00, based on the correlation of the daily power curves of distributed photovoltaics in the same region, select the sample distributed photovoltaics with good quality of the measured data around. Let the number of sample distributed photovoltaics without abnormal data be N. Use the multi-station weighted method (N > 3) or the capacity reduction method (0 < N < 3) to estimate the distributed photovoltaic power generation of the missing or abnormal data. The specific methods are as follows: a. Multi-station weighted method Among them, Figure 4 is the schematic diagram of the principle of the multi-station weighted method.
[0070] Let DG t be the "all-black" distributed photovoltaic to be estimated. Then the power P(t) of the photovoltaic to be estimated at time t is:
[0071] In the formula, P and Q respectively represent the daily power curve and daily power consumption of the distributed photovoltaic.
[0072] b. Capacity reduction method Given the capacity of the distributed photovoltaic to be estimated and the capacity of the adjacent sample distributed photovoltaic, the power of the distributed photovoltaic to be estimated at time t is:
[0073] In the formula, is the capacity of the sample distributed photovoltaic at time t.
[0074] (3) If there is no sample photovoltaic power station with good data quality around the distributed photovoltaic, use the Lagrange interpolation method for completion and correction. The specific method is as follows: Let the 96-point data of the distributed photovoltaic on the same day be as shown in Table 3 below, and the power at time t i is missing, and the power at time t j is abnormal:
[0075] First, find the Lagrange interpolation function P ([[]] t )
[0076] Then, based on the obtained Lagrange interpolation function, complete t. i Missing values and t j Outliers:
[0077] This step considers the daily power operation characteristics of distributed photovoltaic (PV) systems to complete missing outlier data. Data from 0:00-4:00 and 21:00-23:45 are directly set to 0. For other time periods, if there are nearby model distributed PV systems, a multi-site weighted method or capacity reduction algorithm is used to complete missing outlier values. If there are no nearby model distributed PV systems, Lagrange interpolation is used. By fully utilizing the daily power correlation between distributed PV systems in the same region, and based on the PV power data of surrounding model stations, a multi-site weighted method or capacity reduction algorithm is used to complete missing outlier values, thereby improving the data quality of distributed PV-related resource data.
[0078] Figure 6 The structural block diagram of a multi-source heterogeneous distributed photovoltaic data processing device according to an embodiment of this application is shown.
[0079] The multi-source heterogeneous distributed photovoltaic data processing device provided in this application includes the following functional modules: The acquisition module 601 is used to acquire the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regions. The generation module 602 is used to generate a tree-shaped ledger of distributed photovoltaics based on the topological hierarchy and the ledger data of distributed photovoltaics. The determination module 603 is used to determine the distributed photovoltaic power measurement data at each level based on the tree-shaped ledger and the power measurement data of the distributed photovoltaic system. The association module 604 is used to associate the distributed photovoltaic power measurement data of each level with meteorological data based on geographical location; The correction module 605 is used to perform rolling correction on the grid meteorological data in the meteorological data so as to fuse the multi-source meteorological data; The first identification module 606 is used to perform the first identification of the distributed photovoltaic power measurement data of each level based on the distributed photovoltaic operation rules, so as to obtain normal distributed photovoltaic power measurement data. The second identification module 607 is used to perform a second identification on the normal distributed photovoltaic power measurement data based on the outlier detection algorithm to obtain abnormal distributed photovoltaic power measurement data. The correction module 608 is used to complete and correct the abnormal data of the distributed photovoltaic power measurement.
[0080] Optionally, the association module includes: The first submodule is used to load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; The second submodule is used to determine the coordinates of the four corners of the grid point based on the size of the square grid point in the numerical weather forecast and the coordinates of any corner of the grid point; The third submodule is used to determine whether the distributed photovoltaic system is within the grid range; The fourth submodule is used to calculate the distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid range, and to select the nearest meteorological monitoring station as the predictive meteorological data source for the distributed PV system, so as to realize the correlation between distributed PV power measurement data and meteorological data.
[0081] Optionally, the correction module includes: The fifth submodule is used to convert the time interval data monitored by the meteorological monitoring device into data with the same time interval as the numerical weather forecast. The sixth submodule is used to calculate the weighted average error value based on the real-time time-limited data monitored by the meteorological monitoring device on that day. The seventh submodule is used to offset the grid meteorological data in the meteorological data based on the weighted average error value in order to achieve correction.
[0082] Optionally, the correction module includes: The eighth submodule is used to determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; The ninth submodule is used to complete and correct the abnormal data of the distributed photovoltaic power measurement according to the preset completion and correction strategy corresponding to the target time period.
[0083] Optionally, the ninth submodule is specifically used for: If the target time period is a specified nighttime period, the missing abnormal point data in the abnormal data of the distributed photovoltaic power measurement will be set to 0; When the target time period is the first time period, sample distributed photovoltaics that meet the conditions are screened based on the correlation of the daily power curves of distributed photovoltaics in the same region. Based on the number of sample distributed photovoltaic systems that meet the conditions, estimate the distributed photovoltaic power generation power of the distributed photovoltaic power measurement anomaly data that is missing abnormal data. In the absence of a qualified sample distributed photovoltaic system, the Lagrange interpolation method is used to complete and correct the abnormal power measurement data of the distributed photovoltaic system.
[0084] The multi-source heterogeneous distributed photovoltaic (PV) data processing device provided in this application embodiment, firstly, fuses and processes multi-source heterogeneous data such as numerical weather forecasts, meteorological data from meteorological monitoring equipment, and distributed PV power data, improving the accuracy of gridded meteorological data predicted by numerical weather forecasts and achieving full integration of distributed PV-related resource data; secondly, based on distributed PV operation rules and outlier detection algorithms, it identifies and filters outlier anomalies in distributed PV power data in two rounds, making the anomaly data screening criteria more comprehensive and avoiding the adverse effects of missing anomaly data on distributed PV-related application functions; thirdly... By fully utilizing the daily power correlation among distributed photovoltaic (PV) systems in the same region, and using PV power data from surrounding model stations, the missing outliers in distributed PV systems can be supplemented using a multi-site weighting method or a capacity discounting method, thereby improving the data quality of distributed PV-related resource data.
[0085] In the embodiments of this application Figure 6 The multi-source heterogeneous distributed photovoltaic data processing device shown can be installed in a mobile device or a server. The mobile device or server equipped with this device can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application does not specifically limit the specific operating system used.
[0086] The embodiments provided in this application Figure 6 The multi-source heterogeneous distributed photovoltaic data processing device shown can achieve Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0087] Optionally, refer to Figure 7 The present application also provides an electronic device 700, including a processor 701, a memory 702, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the processes executed by the above-mentioned multi-source heterogeneous distributed photovoltaic data processing device and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0088] It should be noted that the electronic device in this application embodiment includes the server described above.
[0089] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0090] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0091] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for processing multi-source heterogeneous distributed photovoltaic data, characterized in that, include: Obtain the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regions; Based on the aforementioned topological hierarchy and the ledger data of distributed photovoltaic power generation, a tree-shaped ledger of distributed photovoltaic power generation is generated. Based on the tree-shaped ledger and the power measurement data of distributed photovoltaics, the power measurement data of distributed photovoltaics at each level are determined; Based on geographical location, the distributed photovoltaic power measurement data at each level are correlated with meteorological data; Rolling correction is performed on the gridded meteorological data in the meteorological data to enable the fusion of multi-source meteorological data; Based on the distributed photovoltaic operation rules, the distributed photovoltaic power measurement data of each level are identified for the first time to obtain normal distributed photovoltaic power measurement data. The normal distributed photovoltaic power measurement data is identified a second time based on the outlier detection algorithm to obtain abnormal distributed photovoltaic power measurement data. The abnormal data in the distributed photovoltaic power measurement were supplemented and corrected. The steps for associating distributed photovoltaic power measurement data at each level with meteorological data based on geographical location include: Load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; Based on the size of the square grid point in the numerical weather forecast and the coordinates of any corner of the grid point, determine the coordinates of the four corners of the grid point; Determine whether distributed photovoltaic systems are within the grid area; The distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid area is calculated, and the nearest meteorological monitoring station is used as the predictive meteorological data source for the distributed PV equipment, so as to realize the correlation between distributed PV power measurement data and meteorological data.
2. The method according to claim 1, characterized in that, The step of performing rolling correction on the gridded meteorological data in the meteorological data to achieve multi-source meteorological data fusion includes: The time interval data monitored by meteorological monitoring devices are converted into data with the same time interval as numerical weather forecasts. The weighted average error value is calculated based on the real-time time-limited data monitored by the meteorological monitoring device on that day. The grid meteorological data in the meteorological data is offset based on the weighted average error value to achieve correction.
3. The method according to claim 1, characterized in that, The steps for supplementing and correcting the abnormal data in the distributed photovoltaic power measurement include: Determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; Based on the preset completion and correction strategy corresponding to the target time period, the abnormal data of the distributed photovoltaic power measurement are completed and corrected.
4. The method according to claim 3, characterized in that, The steps for completing and correcting the abnormal distributed photovoltaic power measurement data according to the preset completion and correction strategy corresponding to the target time period include: If the target time period is a specified nighttime period, the missing abnormal point data in the abnormal data of the distributed photovoltaic power measurement will be set to 0; When the target time period is the first time period, sample distributed photovoltaics that meet the conditions are screened based on the correlation of the daily power curves of distributed photovoltaics in the same region. Based on the number of sample distributed photovoltaic systems that meet the conditions, estimate the distributed photovoltaic power generation power of the distributed photovoltaic power measurement anomaly data that is missing abnormal data. In the absence of a qualified sample distributed photovoltaic system, the Lagrange interpolation method is used to complete and correct the abnormal power measurement data of the distributed photovoltaic system.
5. A multi-source heterogeneous distributed photovoltaic data processing device, characterized in that, include: The acquisition module is used to acquire the topological hierarchy of distributed photovoltaic systems to distribution transformers, feeders, busbars, substations, and regions. The generation module is used to generate a tree-shaped ledger of distributed photovoltaics based on the topological hierarchy and the ledger data of distributed photovoltaics. The determination module is used to determine the distributed photovoltaic power measurement data at each level based on the tree-shaped ledger and the power measurement data of the distributed photovoltaic system. The association module is used to associate the distributed photovoltaic power measurement data of each level with meteorological data based on geographical location; The correction module is used to perform rolling correction on the grid meteorological data in the meteorological data so as to fuse the multi-source meteorological data; The first identification module is used to perform the first identification of the distributed photovoltaic power measurement data of each level based on the distributed photovoltaic operation rules, so as to obtain normal distributed photovoltaic power measurement data. The second identification module is used to perform a second identification on the normal distributed photovoltaic power measurement data based on the outlier detection algorithm to obtain abnormal distributed photovoltaic power measurement data. The correction module is used to complete and correct the abnormal data of the distributed photovoltaic power measurement. The associated module includes: The first submodule is used to load the latitude and longitude coordinates of grid points, meteorological monitoring stations, and distributed photovoltaic systems; The second submodule is used to determine the coordinates of the four corners of the grid point based on the size of the square grid point in the numerical weather forecast and the coordinates of any corner of the grid point; The third submodule is used to determine whether the distributed photovoltaic system is within the grid range; The fourth submodule is used to calculate the distance between each distributed photovoltaic (PV) system and each meteorological monitoring station within the grid range, and to select the nearest meteorological monitoring station as the predictive meteorological data source for the distributed PV system, so as to realize the correlation between distributed PV power measurement data and meteorological data.
6. The apparatus according to claim 5, characterized in that, The correction module includes: The fifth submodule is used to convert the time interval data monitored by the meteorological monitoring device into data with the same time interval as the numerical weather forecast. The sixth submodule is used to calculate the weighted average error value based on the real-time time-limited data monitored by the meteorological monitoring device on that day. The seventh submodule is used to offset the grid meteorological data in the meteorological data based on the weighted average error value in order to achieve correction.
7. The apparatus according to claim 5, characterized in that, The correction module includes: The eighth submodule is used to determine the target time period corresponding to the abnormal data in the distributed photovoltaic power measurement; The ninth submodule is used to complete and correct the abnormal data of the distributed photovoltaic power measurement according to the preset completion and correction strategy corresponding to the target time period.
8. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions are executed by the processor to perform the steps of any one of the multi-source heterogeneous distributed photovoltaic data processing methods according to claims 1-4.