A reference station-based distributed photovoltaic output curve real-time fitting method

CN122456981BActive Publication Date: 2026-09-15上海柒志科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610903911.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-15
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

在实际运行中,非统调光伏数据质量问题具体表现为:台账不统一,同一电站在不同业务系统中名称、容量、坐标信息不一致;数据缺失,通信中断或数据上送异常导致出力曲线出现断点;曲线拉平,采集系统将缺失数据填充为固定值导致出力曲线呈现恒定数值;超装机,台账容量更新不及时导致出力值超过登记容量

Benefits of technology

通过同集群内相邻电站在同一采样时刻的准实时出力值的空间相关性映射实现异常站点的曲线拟合,与现有技术基于气象物理模型或单站历史数据的拟合方式不同,直接利用非统调光伏在已有电能量计量系统中积累的同集群邻居实时数据完成推断。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122456981B_ABST
    Figure CN122456981B_ABST
Patent Text Reader

Abstract

The application discloses a kind of distributed photovoltaic output curve real-time fitting method based on reference station, it is related to photovoltaic output curve fitting technical field, including the following steps: according to the latitude and longitude coordinates of photovoltaic power station and historical output data are carried out spatial clustering division, and the power station similar to space adjacent and output change rule is classified into same photovoltaic fitting cluster;Data quality assessment is carried out based on historical output data, and the station is screened as reference station, and the rest is marked as to be fitted station, with the data complete, integral electric quantity and power energy metering data deviation in the allowable range and the output curve without flattening or jump;The curve fitting of abnormal site is realized by the spatial correlation mapping of the quasi-real-time output value of adjacent power station in the same cluster at the same sampling time, which is different from the fitting mode based on meteorological physical model or single-station historical data of prior art, and directly uses non-uniform photovoltaic to complete inference in existing power energy metering system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of photovoltaic power output curve fitting technology, specifically a real-time fitting method for distributed photovoltaic power output curves based on a reference station. Background Technology

[0002] In power grid dispatching and operation, distributed photovoltaic (PV) systems are categorized into two types based on their management methods: centrally dispatched and non-centrally dispatched. Centrally dispatched PV systems are managed by the power grid dispatch automation system, with their output data transmitted in real-time through the dispatch data network. This results in high-quality data and consistent record keeping. Non-centrally dispatched PV systems are not included in unified dispatch management. Their output data originates from multiple heterogeneous business systems, such as the electricity metering system, marketing business system, and development management system. This leads to issues such as inconsistent record keeping standards, missing active power data, and inconsistent data collection time scales.

[0003] Non-centrally dispatched distributed photovoltaic (PV) systems are numerous, with small individual station capacities and wide spatial distribution, making it impossible to install real-time data acquisition facilities on each one. In actual operation, data quality issues in non-centrally dispatched PV systems manifest in the following ways: inconsistent records, with the same power station showing inconsistent names, capacities, and coordinates in different business systems; missing data, caused by communication interruptions or abnormal data transmission leading to breaks in the output curve; curve flattening, where the acquisition system fills in missing data with fixed values, resulting in a constant output curve; and over-installation, where untimely updates to the record capacity lead to output values ​​exceeding the registered capacity. Summary of the Invention

[0004] There are two main types of existing technologies for fitting distributed photovoltaic power output curves: The first type is based on physical models, which use meteorological cloud maps and irradiance data to build a physical model for photovoltaic power generation and predict output. This type of method requires high-precision meteorological data as input, but non-centrally dispatched distributed photovoltaic systems lack on-site meteorological monitoring, and the spatial density of meteorological stations cannot cover all non-centrally dispatched sites. The physical model also requires accurate component parameters and installation tilt angles, but the ledger information for non-centrally dispatched photovoltaic systems is incomplete, making it impossible to support the parameter configuration of the physical model.

[0005] The second type of modeling is based on statistical interpolation or historical data from individual stations, using linear interpolation or historical output patterns from individual stations to fill in the missing data. Data gaps in non-centrally regulated distributed photovoltaic (PV) systems are persistent; when there are large areas of missing historical data for individual stations, there are insufficient historical samples for modeling. Linear interpolation assumes that output changes are a linear process, which does not match the nonlinear characteristics of actual PV output curves.

[0006] The aforementioned existing technologies do not address the management characteristics of non-centrally dispatched distributed photovoltaic (PV) systems: dispersed data sources, inconsistent record-keeping, and lack of conditions for deploying additional data acquisition facilities. Furthermore, existing technologies do not employ spatial clustering to group non-centrally dispatched PV systems according to their output correlation, nor do they establish a method for repairing the output curves of abnormal sites based on real-time output data from normal sites within the same cluster. To achieve the above objectives, the present invention adopts the following technical solution: Specifically, a real-time fitting method for distributed photovoltaic power output curves based on a benchmark station is proposed, including the following steps: Based on the latitude and longitude coordinates and historical power output data of photovoltaic power stations, spatial clustering is performed to group power stations that are spatially adjacent and have similar power output variation patterns into the same photovoltaic fitting cluster. Based on historical power output data, data quality assessment was conducted. Stations with complete data, deviations between integrated power and electrical energy metering data within the allowable range, and no flattening or abrupt changes in the power output curve were selected as benchmark stations, while the rest were marked as stations to be fitted. At each sampling moment during the fitting period, the near real-time output values ​​of each benchmark station in the cluster where the target power station is located are used as input features to construct a training dataset, train the fitting model, and generate the fitting output curve of the target power station. Physical constraint verification is performed on the fitted power output curve, correcting the portion of the power output value exceeding the rated installed capacity to the rated installed capacity, and correcting the portion of the power output value less than zero to zero.

[0007] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: Curve fitting for abnormal sites is achieved by mapping the spatial correlation of the near real-time output values ​​of adjacent power stations within the same cluster at the same sampling time. Unlike existing technologies that use meteorological and physical models or historical data of a single station for fitting, this method directly utilizes real-time data of neighboring power stations within the same cluster accumulated in the existing power metering system for non-centrally regulated photovoltaic systems to complete the inference.

[0008] Through the dynamic rolling update mechanism of the benchmark station, the data source used to drive the fitting within the cluster is replaced in a timely manner when it becomes abnormal, ensuring that there is a normal data source to support the fitting calculation at each sampling time.

[0009] Physical constraint verification is introduced into the fitting results to ensure that the fitted output curve meets the physical constraints of photovoltaic power generation. Attached Figure Description

[0010] Figure 1 The flowchart shows the fitting process for the real-time fitting method of distributed photovoltaic power output curve. Detailed Implementation

[0011] The technical solution of the present invention will be described below with reference to the accompanying drawings and embodiments.

[0012] Example 1, refer to Figure 1 This embodiment describes the overall process of the real-time fitting method for the output curve.

[0013] Step S1, Ledger Alignment. Obtain ledger information for non-centrally dispatched distributed photovoltaic power stations within the target area from multiple heterogeneous business systems, perform consistency verification on ledgers from different sources, correct conflicting items based on the ledgers registered in the electricity metering system, and establish a unified identity identifier for photovoltaic power stations.

[0014] It should be noted that although the aforementioned non-centrally dispatched photovoltaic systems do not possess the real-time data acquisition capabilities of a dispatch automation system, they still utilize electricity consumption information acquisition terminals (such as concentrators and collectors) within the power metering system to report active power data at fixed acquisition cycles (typically with a time resolution of 15 minutes), forming a near-real-time output sequence with a certain time granularity. The "sampling time" referred to in this invention corresponds to each time breakpoint within the acquisition cycle, such as generating 96 sampling times daily at 15-minute intervals; the "existing data channel" refers to the existing communication link in the power metering system where electricity consumption information acquisition terminals report power data via a dedicated power fiber optic network or a public wireless network. This data channel requires no additional acquisition facilities and can directly utilize the existing metering infrastructure of the power grid company to obtain the near-real-time output values ​​of each power station within the same cluster, providing a data foundation for the spatial correlation mapping of this invention. Step S2, Spatial Clustering. Spatial clustering is performed based on the latitude and longitude coordinates of the photovoltaic power station and its historical power output data.

[0015] Step S3: Data Quality Assessment and Baseline Station Selection. Based on historical power output data, a data quality assessment is conducted. Stations with complete data, integral power consumption and electrical energy metering data deviations within acceptable ranges, and power output curves without flattening or abrupt changes are selected as baseline stations. The remaining stations are marked as stations to be fitted.

[0016] Step S4: Simultaneous real-time data-driven fitting within the same cluster. At each sampling moment during the fitting period, a training dataset is constructed using the near-real-time output values ​​of each benchmark station within the cluster containing the target power station as input features. The fitting model is then trained, and the fitted output curve of the target power station is generated.

[0017] Step S5, Physical constraint verification. Perform physical constraint verification on the fitted power output curve, correcting the portion of the power output value exceeding the rated installed capacity to the rated installed capacity, and correcting the portion of the power output value less than zero to zero.

[0018] Example 2 describes the specific implementation of spatial clustering and baseline station selection.

[0019] In spatial clustering, the latitude and longitude coordinates of the photovoltaic power stations are read, and the Haversine formula is used to calculate the spatial distance matrix between the power stations.

[0020] The Haversine formula calculates the arc distance between two latitude and longitude coordinates based on the geometry of the Earth's sphere. The calculation process is as follows: Let the latitude and longitude coordinates of the two power stations be (lat1, lon1) and (lat2, lon2) respectively. Convert the latitude and longitude to radians, calculate the latitude difference dLat and the longitude difference dLon, and then calculate the intermediate variable a = sin(dLat / 2)^2 + cos(lat1) × cos(lat2) × sin(dLon / 2)^2, and the central angle. The spatial distance d = R × c, where R is the Earth's radius.

[0021] Simultaneously, the Pearson correlation coefficient of historical power output data of photovoltaic power plants is calculated as a measure of the similarity of power output curves. For the historical power output sequences X and Y of two power plants, the Pearson correlation coefficient r = cov(X,Y) / (sigmaX×sigmaY), where cov(X,Y) is the covariance, and sigmaX and sigmaY are the standard deviations of X and Y, respectively. The value of r ranges from [-1,1], and the closer r is to 1, the more synchronous the power output changes of the two power plants are.

[0022] Power plants whose spatial distance is less than a distance threshold and whose output curve Pearson correlation coefficient is greater than a correlation coefficient threshold are grouped into the same photovoltaic fitting cluster, and each cluster is assigned a unique cluster identifier. The dual constraint of spatial distance and output correlation coefficient ensures that power plants that are only spatially adjacent but have unrelated output characteristics will not be grouped into the same cluster, thus ensuring that the output curves of power plants within a cluster have mutually referential patterns of change.

[0023] In one embodiment, the distance threshold is determined based on the statistical results of the spatial distribution density of power plants within the target area. Specifically, the spatial distance between all pairs of power plants is calculated, and the 80th quantile of the distance distribution is used as the initial distance threshold. Then, clustering is performed. If a cluster with fewer than three power plants exists in the clustering results, the distance threshold for the power plants involved in that cluster is increased, and the clustering is repeated until all clusters contain at least three power plants. The correlation coefficient threshold is determined based on the statistical results of the correlation distribution of historical power output data of power plants within the target area. The 60th quantile of the correlation distribution is used as the initial correlation coefficient threshold, and then the mean Pearson correlation coefficient of the power output sequence within each cluster is calculated. If the mean correlation coefficient of a cluster is lower than 0.7, the distance threshold for the power plants involved in that area is reduced, and the clustering is repeated until the mean correlation coefficient of all clusters is not lower than 0.7.

[0024] Under the actual condition of incomplete data, spatial clustering and benchmark station selection are initiated according to the following flexible mechanism: For the calculation of Pearson correlation coefficient required in the spatial clustering stage, data gaps between two power stations are allowed. The correlation coefficient is calculated only using the sampling time when both stations have valid data, requiring an overlap of at least 30 days (equivalent to 2880 sampling points at 15-minute resolution) to participate in the calculation. If the number of benchmark stations that have passed the data quality assessment in a cluster is less than 2, the cluster is merged with the nearest adjacent cluster in terms of spatial distance. The benchmark station selection is then re-executed in the merged cluster until each cluster has at least 2 usable benchmark stations. In extreme cases, if the historical data of all power stations in the target area does not meet the quality assessment standards, a centralized photovoltaic irradiance-output conversion model at the provincial or municipal level is adopted as a cold start scheme. After accumulating at least 7 days of valid operating data, the spatial correlation fitting scheme of this invention is switched to. Example 3 describes the specific implementation of data quality assessment and baseline station selection.

[0025] Data integrity assessment detects the proportion of missing data points in historical output curves. Assuming there should be N sampling points during the assessment period, and M actual valid sampling points, the proportion of missing data points is (N / M) / N. When the proportion of missing data points exceeds the upper limit, the power station's data is deemed incomplete. The upper limit of the missing data proportion is determined based on the historical failure rate statistics of the target area's communication system.

[0026] The data consistency assessment calculates the integral energy of the output curve during the assessment period and compares it with the electrical energy recorded by the electrical energy metering system during the same period. The integral energy Q is obtained by integrating the output curve P(t) over the assessment period [Tstart,Tend]. Q=Integral(Ptdt,Tstart,Tend).

[0027] Under discrete sampling conditions, numerical integration is performed using the trapezoidal rule: Where Deltat is the sampling interval. The integrated energy Q is compared with the energy E recorded by the energy metering system for the same period, and the deviation rate is calculated as |QE| / E. The deviation rate reflects the consistency between the power curve and the energy metering data. When the deviation rate exceeds the upper limit, the data of the power station is determined to be inconsistent. The upper limit of the deviation rate is determined according to the accuracy class of the energy metering device.

[0028] Curve smoothness is assessed by calculating the second-order difference of the force curve to detect flattening and abrupt changes. For a discrete sampling sequence Pi, the first-order difference is DeltaPi = Pi - P{i-1}, and the second-order difference is Delta^2Pi = DeltaPi - DeltaP{i-1} = Pi-2 × P{i-1} + P{i-2}.

[0029] A normal photovoltaic power output curve changes smoothly with illumination conditions, and the absolute value of the second-order difference is relatively small. When the power output curve is filled with fixed values, both the first-order and second-order differences are zero. When a power jump occurs, the absolute value of the second-order difference reaches a maximum. By setting a jump threshold, we can determine whether the absolute value of the second-order difference is abnormal. When the absolute value of the second-order difference exceeds the jump threshold, it is determined that there is flattening or a jump.

[0030] Power plants that pass all three assessments are marked as baseline stations, while power plants that fail any of the assessments are marked as stations to be fitted.

[0031] Example 4 describes the specific implementation of dynamic maintenance of the base station.

[0032] At each sampling moment during the fitting period, the latest output data of each reference station is obtained from the existing data channel, and the following three types of data anomalies are detected: data interruption, that is, no data is sent at the current sampling moment; curve flattening, that is, the output value at the current sampling moment is exactly the same as the output value at the previous two sampling moments; exceeding the installed capacity, that is, the output value at the current sampling moment exceeds the rated installed capacity of the power station.

[0033] If any of the above-mentioned data anomalies are found in the output data of a base station at the current sampling time, the data of that base station will be removed from the training samples at the current sampling time. Within its cluster, base stations will be re-selected based on their comprehensive data quality assessment score from the previous day. The comprehensive score is calculated as follows: The overall score is calculated as follows: w1 / (proportion of missing points + epsilon) + w2 / (bias rate + epsilon) + w3 / (smoothness index + epsilon), where w1, w2, and w3 are the weighting coefficients for the three evaluations, and epsilon is a small constant to prevent division by zero. The station with the highest score is selected as the new baseline station. Starting from the next sampling time, the new baseline station's output data is included in the training samples, and it takes over the power output data provision from the original baseline station.

[0034] When there are no power stations in the cluster that meet the criteria for the baseline station selection, the cluster is merged with the nearest adjacent cluster in terms of spatial distance. After the merger, the power stations of the two clusters together form a new fitting cluster, and their respective output data are used to construct the training dataset.

[0035] Example 5 describes the specific implementation of a real-time data-driven fitting strategy within the same cluster.

[0036] The fitting model takes the combined output values ​​of each reference station within the same cluster at the same sampling time as input, and the measured output value of the target power station at the corresponding time during historical normal periods as output, learning the spatial correlation mapping relationship of the output curves within the cluster. The essence of this mapping relationship is that, under normal operating conditions, the output curves of photovoltaic power stations within the same cluster change synchronously with illumination conditions, but due to the influence of installed capacity, orientation, tilt angle, and local shading factors, there is a nonlinear correspondence between the output values ​​of photovoltaic power stations. The fitting model learns this nonlinear correspondence through a large amount of simultaneous output data under normal operating conditions.

[0037] Let the output value of the target power station be y, and the output values ​​of n reference stations within the same cluster at the same sampling time be x1, x2, ..., xn, respectively. Then, the fitting model learns the mapping function f, such that y = f(x1, x2, ..., xn, C, T), where C is the installed capacity of the target power station and T is the time feature vector. When the data of the target power station at the current sampling time is abnormal, the near real-time output values ​​x1', x2', ..., xn' of each reference station within the same cluster at that sampling time are substituted into the mapping function to obtain the inferred output value of the target power station y' = f(x1', x2', ..., xn', C, T).

[0038] This inference mechanism is a spatial mapping, fundamentally different from the temporal-dimensional prediction of existing technologies: prediction methods use historical data to infer future output values, while this invention uses the near-real-time output values ​​of neighboring power plants in the same cluster at the same moment to infer the output value of the target power plant at the same moment. When the target power plant's current sampling time data is abnormal, this invention does not require historical data or meteorological data, but only the near-real-time output values ​​of neighboring power plants in the same cluster at that sampling time to complete the inference.

[0039] Example 6 describes the specific implementation of constructing the training dataset.

[0040] The input features of the training dataset consist of three parts: the near real-time power output values ​​x1, x2, ..., xn of each reference station within the same cluster at the same sampling time; the installed capacity C of the target power station; and time features T, including at least one of month code, time code, and season code. The month code uses integers from 1 to 12, the time code is determined according to the sampling interval, and the season code aggregates months into four categories: spring, summer, autumn, and winter.

[0041] The output values ​​in the input features are processed using Min-Max normalization to eliminate the dimensional effects caused by differences in installed capacity among different power plants. The formula for Min-Max normalization is: xnorm = (x - xmin) / (xmax - xmin), where x is the original output value, and xmin and xmax are the minimum and maximum output values ​​of all benchmark stations in the cluster within the training data window, respectively. After normalization, the output values ​​are linearly mapped to the [0,1] interval, enabling the model to learn the relative variation patterns between output values.

[0042] The output label of the training dataset is the measured power output value of the target power plant at the corresponding time during the historical normal period. During training, the model learns to predict the power output value of the target power plant under given input features. After training, when the data of the target power plant at the current sampling time is abnormal, the near real-time power output value of the same cluster benchmark station at that sampling time, the installed capacity of the target power plant, and the time features are input into the trained model, and the inferred power output value of the target power plant is output.

[0043] Example 7 describes the specific implementation of the multi-model fusion strategy.

[0044] Gradient boosting tree model, extreme gradient boosting model, and long short-term memory network model were selected as base learners. The gradient boosting tree model, by iteratively training a decision tree to fit pseudo-residuals, has a natural ability to handle categorical features and is suitable for processing discrete features such as power plant numbers and seasonal codes. The extreme gradient boosting model has advantages in second-order gradient optimization and regularization, achieving high fitting accuracy and fast training speed. The long short-term memory network model controls information flow through a gating mechanism, excelling at capturing long-term dependencies in time series data and is suitable for modeling the temporal variation patterns of power output curves. These three models complement each other in feature processing, optimization strategies, and temporal modeling.

[0045] Three base learners were independently trained on the training dataset to obtain independent predictions of the target power plant's output value for each base learner. The fusion weights were calculated based on the mean absolute percentage error (MAPEi) of each base learner on historical validation data. Let the mean absolute percentage error of the i-th base learner be MAPEi, then its fusion weight Wi = (1 / MAPEi) / Sum(1 / MAPEj), where j iterates through all base learners. The formula for calculating the mean absolute percentage error is: MAPE = (1 / N) × Sum(|ytrue - ypred| / |ytrue|) × 100%, where N is the number of validation samples, ytrue is the measured output force value, and ypred is the model prediction value. The weighted fusion result is the weighted sum of the independent prediction results of each base learner according to the fusion weights.

[0046] Model iterative optimization comprises two parts: recalculation of fusion weights and retraining on training data. The fusion weights of each base learner are recalculated at fixed time intervals to ensure they reflect the actual fitting ability of each base learner under the current seasonal conditions. Full retraining of the training dataset is performed at fixed time intervals, discarding data from the earliest training data window and incorporating data from the latest training data window during retraining. This allows the model parameters to adapt to output characteristic drift caused by seasonal changes and equipment aging. The retraining interval is shortened in seasons with rapid output characteristic changes and lengthened in seasons with slow changes.

[0047] Example 8 describes the specific implementation of physical constraint verification.

[0048] Over-installed capacity constraint verification: For each sampling point Pfitted(t) in the fitted power output curve, if Pfitted(t) is greater than the rated installed capacity Prated of the power station, then Pfitted(t) is corrected to Prated. The triggering of over-installed capacity indicates that the fitted model overestimates the power output capacity of the target power station.

[0049] Negative value constraint verification: For each sampling point Pfitted(t) in the fitted output curve, if Pfitted(t) is less than 0, then Pfitted(t) is corrected to 0. The triggering of a negative value indicates that the fitted model underestimates the output capacity of the target power plant.

[0050] The triggering frequency of the two constraint checks serves as a model performance monitoring indicator. When the triggering frequency exceeds the monitoring threshold, the retraining process of the fitted model is triggered.

[0051] Non-centrally dispatched distributed photovoltaic: refers to distributed photovoltaic power stations that are not included in the unified management of the power grid dispatch automation system. Their output data comes from multiple heterogeneous business systems such as the power metering system and the marketing business system, and is not transmitted through the dispatch data network.

[0052] Photovoltaic fitting cluster: refers to a set of power stations obtained by spatial clustering. The photovoltaic power stations in the set are spatially close and their historical output curves are correlated. The output curves of the power stations in the set can be referenced by each other.

[0053] Baseline station: refers to a power station whose data quality meets the standards after being screened through data quality assessment. Its output data is used to build a training dataset to drive the training and inference of the fitting model.

[0054] Stations to be fitted: These are power plants that have not passed the data quality assessment. Their output data is missing, biased, or abnormal, and their output curves need to be inferred through a fitting model.

[0055] Spatial correlation mapping relationship: refers to the nonlinear correspondence between the output values ​​of photovoltaic power plants within the same photovoltaic fitting cluster. This relationship is learned by the fitting model through the synchronous output data of photovoltaic power plants during historical normal periods within the cluster.

[0056] Simultaneous inference: When the data of the target power station at the current sampling time is abnormal, the near real-time output value of the benchmark station in the same cluster at that sampling time is input into the trained fitting model, and the model outputs the inferred output value of the target power station at that sampling time. This process is a mapping in the spatial dimension rather than a prediction in the temporal dimension.

[0057] Integral power generation: refers to the definite integral of the power curve over time during the assessment period, representing the estimated cumulative power generation during the assessment period.

[0058] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for real-time fitting of distributed photovoltaic power output curves based on a reference station, characterized in that, Includes the following steps: Based on the latitude and longitude coordinates and historical power output data of each power station, spatial clustering is performed to group power stations that are spatially adjacent and have similar power output variation patterns into the same photovoltaic fitting cluster. Data quality assessment is conducted based on historical output data. Stations with complete data, whose integral power output and electrical energy metering data deviations are within acceptable limits, and whose output curves show no flattening or abrupt changes are selected as benchmark stations. The remaining stations are marked as stations to be fitted. The data quality assessment includes: data integrity assessment, checking whether the proportion of missing data points in the historical output curve exceeds the upper limit for missing proportions; data consistency assessment, calculating the integral power output of the output curve during the assessment period, comparing the integral power output with the electrical energy recorded by the electrical energy metering system during the same period, calculating the deviation rate, and determining whether the deviation rate exceeds the upper limit for deviation rate; and curve smoothness assessment, calculating the second difference of the output curve, and determining whether the output curve shows flattening or abrupt changes based on whether the absolute value of the second difference exceeds the abrupt change threshold. At each sampling moment during the fitting period, the real-time output value of each benchmark station in the cluster where the target power station is located is used as the input feature to construct a training dataset, train the fitting model, and generate the fitting output curve of the target power station. The dynamic maintenance steps for benchmark stations are as follows: At each sampling moment during the fitting period, the latest output data of each benchmark station is obtained from the existing data channels to detect three types of data anomalies: data interruption, curve flattening, or exceeding the installed capacity. When the output data of a benchmark station at the current sampling moment is abnormal, the data of that benchmark station is removed from the training sample at the current sampling moment, and the benchmark stations are re-selected within its cluster according to the comprehensive score of the data quality assessment of the previous day. The output data of the newly selected benchmark stations are included in the training sample from the next sampling moment, and the new benchmark stations take over the output data provided by the original benchmark stations. When there are no stations in the cluster that meet the benchmark station selection criteria, the cluster is merged with the nearest adjacent cluster in terms of spatial distance. After the merger, the power stations of the two clusters together form a new fitting cluster, and their respective output data jointly participate in the construction of the training dataset. Physical constraint verification is performed on the fitted power output curve, correcting the portion of the power output value exceeding the rated installed capacity to the rated installed capacity, and correcting the portion of the power output value less than zero to zero.

2. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 1, characterized in that, The spatial clustering partitioning includes: Calculate the distance matrix between each power station using spatial distance as a similarity metric; The correlation coefficient of historical power output data is used as a measure of the similarity of power output curves; Power plants whose spatial distance is less than the distance threshold and whose power output curve correlation coefficient is greater than the correlation coefficient threshold are grouped into the same photovoltaic fitting cluster.

3. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 2, characterized in that, The methods for determining the distance threshold and the correlation coefficient threshold include: The distance threshold is determined based on the spatial distribution density statistics of power stations within the target area, so that each cluster contains no less than the preset minimum number of power stations. The correlation coefficient threshold is determined based on the statistical results of the correlation distribution of historical power output data of power plants in the target area, so that the average correlation coefficient of the power output curve of each power plant in the cluster is not lower than the preset minimum level.

4. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 1, characterized in that, The training mechanism of the fitted model includes: The fitting model takes the combination of power output values ​​of each reference station in the same cluster at the same sampling time as input and the measured power output value of the target power station at the corresponding time during the historical normal period as output, and learns the spatial correlation mapping relationship of the power output curve within the cluster. When the output data of the target power station at the current sampling time is abnormal, the real-time output values ​​of each benchmark station in the cluster at that sampling time are input into the trained fitting model, and the fitting model outputs the inferred output value of the target power station at that sampling time.

5. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 4, characterized in that, The construction of the training dataset also includes: The installed capacity and time characteristics of the target power plant are included in the input features, wherein the time characteristics include at least one of month, time of day, and season; The output values ​​in the input features are normalized using the Min-Max normalization method. The normalization scaling factor is the difference between the maximum and minimum output values ​​of all base stations in the cluster. The measured power output value of the target power plant at the corresponding moment during the historical normal period is used as the output label of the training sample.

6. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 1, characterized in that, The fitting model is constructed using a multi-model fusion strategy, including: Gradient boosting tree model, extreme gradient boosting model and long short-term memory network model are selected as base learners. The three base learners are complementary in feature processing, gradient optimization and temporal modeling, respectively. Each base learner is trained independently on the training dataset to obtain independent prediction results of the target power plant output value by each base learner. The fusion weights of each base learner are calculated based on the average absolute percentage error of each base learner on historical validation data. The independent prediction results of each base learner are then weighted and summed according to the fusion weights to obtain the fusion prediction result.

7. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 6, characterized in that, The fusion weight is calculated as follows: The fusion weight of each base learner is equal to the ratio of the reciprocal of the mean absolute percentage error of that base learner to the sum of the reciprocals of the mean absolute percentage errors of all base learners; The mean absolute percentage error is the average percentage of the absolute deviation between the inferred force value and the measured force value at each sampling time on the verification data, relative to the measured force value.

8. The method for real-time fitting of distributed photovoltaic power output curves based on a reference station according to claim 7, characterized in that, It also includes model iterative optimization steps: The fusion weights of each base learner are recalculated at fixed time intervals so that the fusion weights reflect the actual fitting ability of each base learner under the current seasonal conditions. The training dataset is fully retrained at fixed time intervals. During retraining, the data in the earliest training data window is discarded and the data in the latest training data window is included, so that the model parameters can adapt to the output characteristics drift caused by seasonal changes and equipment aging.

Citation Information

Patent Citations

  • Regional distributed photovoltaic power prediction method and device, computer equipment and storage medium

    CN116316542A

  • Distributed photovoltaic power data reconstruction method based on spatial correlation

    CN117235450A

  • Method and device for calculating generating capacity of photovoltaic power station, electronic equipment and storage medium

    CN120408015A

  • Hybrid photovoltaic power prediction method and system based on multi-source data fusion

    US20220373984A1