A photovoltaic power generation power prediction method, device and computer readable storage medium

By calculating the correlation coefficient between meteorological parameters and photovoltaic power generation, clustering and weighting are performed. Combined with Bayesian linear regression and Gaussian process regression models, the problem that the influence of meteorological parameters was not considered in the existing technology is solved, and the accuracy of photovoltaic power generation prediction is improved.

CN119765339BActive Publication Date: 2026-01-23STATE GRID SHANXI MARKETING SERVICE CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510261814.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2026-01-23
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

Existing photovoltaic power generation forecasting methods ignore the impact of different meteorological parameters on photovoltaic power generation, resulting in low accuracy of forecast results.

Method used

By calculating the correlation coefficient between meteorological parameters and photovoltaic power generation, clustering and weighting are performed. A probability distribution model for predicting photovoltaic power generation is constructed by combining Bayesian linear regression and Gaussian process regression models, and the model parameters are optimized using a genetic algorithm.

Benefits of technology

It improves the accuracy of photovoltaic power generation forecasting, reduces the impact of mismatched meteorological parameters on forecast results, and provides more accurate forecast results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119765339B_ABST
    Figure CN119765339B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of photovoltaic power generation, and relates to a photovoltaic power generation power prediction method, device and computer readable storage medium; meteorological parameters and photovoltaic power generation power at multiple sampling time points are collected to obtain photovoltaic power generation data sequences at multiple sampling time points, and the correlation coefficient of each meteorological parameter and photovoltaic power generation power is calculated; K photovoltaic power generation data sequences are taken as clustering centers, all photovoltaic power generation data sequences are clustered to obtain K initial photovoltaic power generation data clusters; the local outlier factor and weight of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster are calculated; based on the weight of all photovoltaic power generation data sequences and the posterior distribution of a Bayesian linear regression model, a photovoltaic power generation power prediction probability distribution model is constructed; the parameters of the photovoltaic power generation power prediction probability distribution model are optimized by using the photovoltaic power generation data sequences, a target photovoltaic power generation power prediction model is obtained, and the accuracy of the photovoltaic power generation power prediction result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power generation technology, and in particular to a photovoltaic power generation prediction method, apparatus, and computer-readable storage medium. Background Technology

[0002] With the gradual reduction of global reliance on traditional fossil fuels and increasing attention to environmental pollution, clean energy is rapidly emerging as a key pillar of the energy structure. Solar energy, as an easily accessible and environmentally friendly renewable energy source, has gained widespread recognition and application in the power industry. However, the output power of photovoltaic (PV) power generation is highly susceptible to weather conditions, exhibiting significant intermittency and volatility. This instability poses challenges to the stability and manageability of the power system when PV power is integrated into the grid on a large scale. Therefore, strengthening the accurate prediction of PV power output is of great significance to ensure the stable and safe operation of the power system.

[0003] Traditional deep learning-based photovoltaic (PV) power prediction methods train neural network models using historical PV power data, enabling the models to learn patterns in historical PV power output and predict future output. However, PV power output is often influenced by meteorological factors, resulting in significant differences in the PV power output curve under different meteorological conditions. For this reason, existing PV power prediction methods often collect historical PV power data and corresponding meteorological parameters as input to the model during training. This allows the model to learn the correlation between PV power output and meteorological conditions, thus considering the impact of meteorological factors in PV power prediction and making the prediction results more accurate. However, since different meteorological parameters, such as temperature, humidity, and radiation intensity, have varying degrees of influence on PV power output, existing methods that train models based on meteorological parameters and historical PV power data ignore this factor. They fail to capture the degree of influence of each meteorological parameter on PV power output, leading to significant errors in the final prediction results.

[0004] In summary, existing photovoltaic power generation prediction methods ignore the impact of different meteorological parameters on photovoltaic power generation, resulting in photovoltaic power generation prediction models that cannot capture the relationship between different meteorological parameters and photovoltaic power generation, thus leading to low accuracy of prediction results. Summary of the Invention

[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem that the photovoltaic power generation prediction method in the prior art ignores the influence of different meteorological parameters on photovoltaic power generation, which leads to the photovoltaic power generation prediction model being unable to capture the relationship between different meteorological parameters and photovoltaic power generation, resulting in low accuracy of the prediction results.

[0006] To address the aforementioned technical problems, this invention provides a method for predicting photovoltaic power generation, comprising:

[0007] Meteorological parameters and photovoltaic power generation at multiple sampling times were collected to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, the correlation coefficient between each meteorological parameter and photovoltaic power generation was calculated.

[0008] K photovoltaic power generation data sequences are used as cluster centers. Based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, the Euclidean distance between each photovoltaic power generation data sequence and each cluster center is calculated. Thus, all photovoltaic power generation data sequences are clustered to obtain K initial photovoltaic power generation data clusters.

[0009] Calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence;

[0010] Based on the weights of all photovoltaic power generation data sequences, a diagonal weight matrix is ​​constructed; based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model, a probability distribution model for photovoltaic power generation prediction is constructed.

[0011] The multiple photovoltaic power generation data sequences are input into the photovoltaic power generation prediction probability distribution model, and the parameters of the photovoltaic power generation prediction probability distribution model are optimized to obtain the target photovoltaic power generation prediction model.

[0012] Preferably, the formula for calculating the correlation coefficient between each meteorological parameter and photovoltaic power generation is as follows:

[0013] ,

[0014] in, Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first In the photovoltaic power generation data sequence, the first Various meteorological parameters; This represents the first [number]th ... The average value of various meteorological parameters; Indicates the first Photovoltaic power generation in a photovoltaic power generation data sequence; This represents the average photovoltaic power generation across all photovoltaic power generation data series. Indicates the number of photovoltaic power generation data sequences; , Indicates the types of meteorological parameters.

[0015] Preferably, the formula for calculating the Euclidean distance between each photovoltaic power generation data sequence and each cluster center is expressed as:

[0016] ,

[0017] in, Indicates the first A photovoltaic power generation data sequence With the kth cluster center Euclidean distance; Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first A photovoltaic power generation data sequence The Middle Various meteorological parameters; Represents the k-th cluster center The Middle Various meteorological parameters; , Indicates the types of meteorological parameters.

[0018] Preferably, after obtaining K initial photovoltaic power generation data clusters, the process further includes:

[0019] Based on the class centers of the K initial photovoltaic power generation data clusters, K new cluster centers are obtained;

[0020] Based on the K new cluster centers, all photovoltaic power generation data sequences are re-clustered to obtain K new photovoltaic power generation data clusters, and it is determined whether the K new photovoltaic power generation data clusters are the same as the K initial photovoltaic power generation data clusters.

[0021] If the K new photovoltaic power generation data clusters are different from the K initial photovoltaic power generation data clusters, then the class centers of the K new photovoltaic power generation data clusters are used as the K new cluster centers, and all photovoltaic power generation data sequences are re-clustered until the K photovoltaic power generation data clusters obtained by clustering are the same as the K new photovoltaic power generation data clusters.

[0022] Preferably, the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster is calculated, and the weight of each photovoltaic power generation data sequence is calculated based on the local outlier factor of each photovoltaic power generation data sequence, including:

[0023] For each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, the k-th photovoltaic power generation data sequence that is closest to the photovoltaic power generation data sequence in the initial photovoltaic power generation data cluster is taken as the target photovoltaic power generation data sequence, and the distance between the photovoltaic power generation data sequence and the target photovoltaic power generation data sequence is taken as the K-nearest neighbor distance of the photovoltaic power generation data sequence;

[0024] Obtain photovoltaic power generation data sequences from the initial photovoltaic power generation data cluster whose distance to the photovoltaic power generation data sequence is less than the K-nearest neighbor distance, and obtain a set of photovoltaic power generation data sequences;

[0025] The maximum value between the K-nearest neighbor distance of the photovoltaic power generation data sequence and the K-nearest neighbor distance of the target photovoltaic power generation data sequence is taken as the reachability distance of the photovoltaic power generation data sequence;

[0026] Based on the reachability distance of the photovoltaic power generation data sequence and the set of photovoltaic power generation data sequences, calculate the local reachability density of the photovoltaic power generation data sequence;

[0027] The local outlier factor of the photovoltaic power generation data sequence is calculated based on the local reachability density of the photovoltaic power generation data sequence.

[0028] The local outlier factors of the photovoltaic power generation data sequence are mapped based on preset coefficients to obtain the weights of the photovoltaic power generation data sequence.

[0029] Preferably, the formula for calculating the local reachability density of the photovoltaic power generation data sequence is:

[0030] ,

[0031] in, Represents photovoltaic power generation data sequence Locally achievable density; Represents a photovoltaic power generation data sequence The k-neighborhood, i.e., the photovoltaic power generation data sequence Centered on a circle, using photovoltaic power generation data sequences The neighborhood with a radius equal to the K-nearest neighbor distance; Represents photovoltaic power generation data sequence The reachable distance;

[0032] The formula for calculating the local outlier factor of a photovoltaic power generation data sequence is as follows:

[0033] ,

[0034] in, This represents the local outlier factor in a photovoltaic power generation data sequence. Represents photovoltaic power generation data sequence Locally achievable density;

[0035] The formula for calculating the weights of a photovoltaic power generation data sequence is as follows:

[0036] ,

[0037] in, Indicates the first Weights of individual photovoltaic power generation data sequences; This indicates the preset coefficient.

[0038] Preferably, the photovoltaic power generation prediction probability distribution model is expressed as:

[0039] ,

[0040] ,

[0041] in, This represents the probability distribution model for predicting photovoltaic power generation. This represents the predicted photovoltaic power generation value output by the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters input into the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters in the photovoltaic power generation data series; This represents the photovoltaic power generation in the photovoltaic power generation data sequence; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; This represents the diagonal weight matrix; Noise indicating photovoltaic power generation output; Represents the identity matrix.

[0042] Preferably, after inputting the multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, the parameters of the photovoltaic power generation prediction probability distribution model are optimized using a genetic algorithm.

[0043] The present invention also provides a photovoltaic power generation prediction device, comprising:

[0044] The data acquisition and correlation coefficient calculation module is used to collect meteorological parameters and photovoltaic power generation at multiple sampling times to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, the correlation coefficient between each meteorological parameter and photovoltaic power generation is calculated.

[0045] The clustering module is used to take K photovoltaic power generation data sequences as cluster centers and calculate the Euclidean distance between each photovoltaic power generation data sequence and each cluster center based on the correlation coefficient between each meteorological parameter and photovoltaic power generation. This allows for the clustering of all photovoltaic power generation data sequences to obtain K initial photovoltaic power generation data clusters.

[0046] The weight calculation module is used to calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence.

[0047] The model building module is used to construct a diagonal weight matrix based on the weights of all photovoltaic power generation data sequences; and to construct a photovoltaic power generation prediction probability distribution model based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model.

[0048] The parameter optimization and model acquisition module is used to input the multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, optimize the parameters of the photovoltaic power generation prediction probability distribution model, and obtain the target photovoltaic power generation prediction model.

[0049] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the photovoltaic power generation prediction method described above.

[0050] The photovoltaic power generation prediction method provided in this application first calculates the correlation coefficient between photovoltaic power generation and various meteorological parameters based on photovoltaic power generation at multiple sampling times. By calculating the correlation coefficient, the linear relationship between different meteorological parameters and photovoltaic power generation is explored, revealing the degree of influence of each meteorological parameter on photovoltaic power generation. Based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, all photovoltaic power generation data sequences are clustered to obtain K initial photovoltaic power generation data clusters. The feature weight K-means clustering method is then used to cluster the photovoltaic power generation data sequences, grouping those belonging to the same meteorological type into one data cluster. Due to the influence of data acquisition equipment or environmental factors, some photovoltaic power generation data may exhibit anomalies, meaning that some photovoltaic power generation and its meteorological parameters do not completely correspond or couple. Therefore, the local outlier factor algorithm is used to calculate the local outlier factor of the photovoltaic power generation data sequences in each meteorological type data cluster to determine the photovoltaic power generation data sequence's local outlier factor. The degree of coupling between photovoltaic (PV) power generation and meteorological parameters in the PV power generation data sequence is used to determine whether the PV power generation data sequence is abnormal and assign corresponding weights to it. Based on the weights of all PV power generation data sequences and the posterior distribution of the Bayesian linear regression model, a PV power generation prediction probability distribution model is constructed. Finally, the parameters of the PV power generation prediction probability distribution model are optimized using multiple PV power generation data sequences to obtain the target PV power generation prediction model. This application fully considers the impact of different meteorological parameters on PV power generation. By clustering PV power generation data sequences corresponding to different meteorological parameter types, and combining the local outlier factor algorithm with the Gaussian process regression model, each PV power generation data sequence is assigned a corresponding weight value. This allows the model to learn the correlation between different meteorological parameters and PV power generation, and also reduces the impact of abnormal data where PV power generation does not match meteorological parameters on the model's prediction results, greatly improving the accuracy of PV power generation prediction results. Attached Figure Description

[0051] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:

[0052] Figure 1 Here is a flowchart of the photovoltaic power generation prediction method provided in this application;

[0053] Figure 2 This is a schematic diagram of the photovoltaic power generation prediction device provided in this application. Detailed Implementation

[0054] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0055] Please see Figure 1 , Figure 1 The flowchart of the photovoltaic power generation prediction method provided in this application is as follows:

[0056] S10: Collect meteorological parameters and photovoltaic power generation at multiple sampling times to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, calculate the correlation coefficient between each meteorological parameter and photovoltaic power generation.

[0057] Specifically, meteorological parameters include temperature, humidity, radiation intensity, and brightness;

[0058] S20: Using K photovoltaic power generation data sequences as cluster centers, and based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, calculate the Euclidean distance between each photovoltaic power generation data sequence and each cluster center, thereby clustering all photovoltaic power generation data sequences to obtain K initial photovoltaic power generation data clusters;

[0059] S30: Calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence;

[0060] S40: Construct a diagonal weight matrix based on the weights of all photovoltaic power generation data sequences; construct a probability distribution model for photovoltaic power generation prediction based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model.

[0061] S50: Input multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, optimize the parameters of the photovoltaic power generation prediction probability distribution model, and obtain the target photovoltaic power generation prediction model.

[0062] The photovoltaic power generation prediction method provided in this application first calculates the correlation coefficient between photovoltaic power generation and various meteorological parameters based on photovoltaic power generation at multiple sampling times. By calculating the correlation coefficient, the linear relationship between different meteorological parameters and photovoltaic power generation is explored, revealing the degree of influence of each meteorological parameter on photovoltaic power generation. Based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, all photovoltaic power generation data sequences are clustered to obtain K initial photovoltaic power generation data clusters. The feature weight K-means clustering method is then used to cluster the photovoltaic power generation data sequences, grouping those belonging to the same meteorological type into one data cluster. Due to the influence of data acquisition equipment or environmental factors, some photovoltaic power generation data may exhibit anomalies, meaning that some photovoltaic power generation and its meteorological parameters do not completely correspond or couple. Therefore, the local outlier factor algorithm is used to calculate the local outlier factor of the photovoltaic power generation data sequences in each meteorological type data cluster to determine the photovoltaic power generation data sequence's local outlier factor. The degree of coupling between photovoltaic (PV) power generation and meteorological parameters in the PV power generation data sequence is used to determine whether the PV power generation data sequence is abnormal and assign corresponding weights to it. Based on the weights of all PV power generation data sequences and the posterior distribution of the Bayesian linear regression model, a PV power generation prediction probability distribution model is constructed. Finally, the parameters of the PV power generation prediction probability distribution model are optimized using multiple PV power generation data sequences to obtain the target PV power generation prediction model. This application fully considers the impact of different meteorological parameters on PV power generation. By clustering PV power generation data sequences corresponding to different meteorological parameter types, and combining the local outlier factor algorithm with the Gaussian process regression model, each PV power generation data sequence is assigned a corresponding weight value. This allows the model to learn the correlation between different meteorological parameters and PV power generation, and also reduces the impact of abnormal data where PV power generation does not match meteorological parameters on the model's prediction results, greatly improving the accuracy of PV power generation prediction results.

[0063] Since the power generation of photovoltaic panels is the result of the combined influence of various factors during operation, it is necessary to consider the influence of multiple meteorological factors. At the same time, the shape of the photovoltaic power curve varies greatly under different meteorological parameters, and the prediction accuracy of the model trained using data under different parameters is also different. Therefore, the embodiments of this application calculate the correlation coefficient between various meteorological parameters and photovoltaic power generation.

[0064] Specifically, this embodiment quantifies the linear relationship between different meteorological parameters and photovoltaic power generation through the Pearson correlation coefficient, and its specific calculation formula is expressed as follows:

[0065] ,

[0066] in, Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first In the photovoltaic power generation data sequence, the first Various meteorological parameters; This represents the first [number]th ... The average value of various meteorological parameters; Indicates the first Photovoltaic power generation in a photovoltaic power generation data sequence; This represents the average photovoltaic power generation across all photovoltaic power generation data series. Indicates the number of photovoltaic power generation data sequences; , Indicates the types of meteorological parameters.

[0067] Furthermore, after calculating the correlation coefficients between various meteorological parameters and photovoltaic power generation, the power generation data sequence is classified by characteristic weight K-means clustering. This aims to identify different local meteorological types, more accurately capture the impact of meteorological factors on photovoltaic power generation, and thus provide a more reliable basis for the prediction of photovoltaic power generation.

[0068] K-means clustering is a clustering algorithm based on the partitioning of a sample set. Its basic idea is to divide data points into K clusters such that the sum of the distances between each data point and the center of its respective cluster is minimized. This method can efficiently process large datasets. This application combines the results of correlation analysis with traditional K-means clustering to form feature-weighted K-means clustering. It adds the concept of feature weights to the traditional K-means clustering algorithm to handle situations where different features may have different importance to the clustering results.

[0069] Specifically, this application generates a weighted Euclidean distance by weighting the traditional Euclidean distance, which can more accurately classify weather types. Optionally, the features can be standardized before weighting to eliminate the influence of different units and numerical ranges.

[0070] Specifically, the formula for calculating the Euclidean distance between each photovoltaic power generation data sequence and each cluster center is as follows:

[0071] ,

[0072] in, Indicates the first A photovoltaic power generation data sequence With the kth cluster center Euclidean distance; Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first A photovoltaic power generation data sequence The Middle Various meteorological parameters; Represents the k-th cluster center The Middle Various meteorological parameters; , Indicates the types of meteorological parameters.

[0073] Optionally, in some embodiments of this application, in order to make the clustering results more accurate, after obtaining K initial photovoltaic power generation data clusters, the method further includes:

[0074] Based on the class centers of the K initial photovoltaic power generation data clusters, K new cluster centers are obtained;

[0075] Based on the K new cluster centers, all photovoltaic power generation data sequences are re-clustered to obtain K new photovoltaic power generation data clusters, and it is determined whether the K new photovoltaic power generation data clusters are the same as the K initial photovoltaic power generation data clusters.

[0076] If the K new photovoltaic power generation data clusters are different from the K initial photovoltaic power generation data clusters, then the class centers of the K new photovoltaic power generation data clusters are used as the K new cluster centers, and all photovoltaic power generation data sequences are re-clustered until the K photovoltaic power generation data clusters obtained by clustering are the same as the K new photovoltaic power generation data clusters.

[0077] During photovoltaic power generation, data acquisition, and transmission, human error, environmental factors such as electromagnetic interference, and excessively high temperatures can all affect the normal operation of data acquisition equipment. Communication delays between the data acquisition equipment and the data center can lead to data lag, and photovoltaic power generation and meteorological data cannot be fully correlated or coupled. Therefore, it is necessary to process abnormal photovoltaic power generation data sequences. However, the reduction of the dataset will simultaneously lead to a reduction in model training. Therefore, this application adopts a weighted approach to reduce the weight of abnormal photovoltaic power generation data sequences, thereby improving prediction accuracy without affecting the amount of data.

[0078] Specifically, the Local Outlier Factor (LOF) algorithm is an unsupervised anomaly detection algorithm based on data point density. It identifies outliers by comparing the local density differences between a data point and its neighbors, and considers samples with a local density significantly lower than its nearest neighbors as anomalies. The algorithm calculates a local outlier factor for each data point and observes whether the LOF value deviates significantly from 1 to determine whether a point is an anomaly. If the LOF value is much higher than 1, the data point is likely to be an anomaly. If the LOF value is close to 1, the data point is usually considered a normal data point.

[0079] As can be seen from the principle of the LOF algorithm, the key to detecting whether data values ​​are abnormal is to calculate the LOF value. Therefore, the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster is calculated. Based on the local outlier factor of each photovoltaic power generation data sequence, the weight of each photovoltaic power generation data sequence is calculated, including:

[0080] For each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, the k-th photovoltaic power generation data sequence that is closest to the photovoltaic power generation data sequence in the initial photovoltaic power generation data cluster is taken as the target photovoltaic power generation data sequence, and the distance between the photovoltaic power generation data sequence and the target photovoltaic power generation data sequence is taken as the K-nearest neighbor distance of the photovoltaic power generation data sequence;

[0081] Obtain photovoltaic power generation data sequences from the initial photovoltaic power generation data cluster whose distance to the photovoltaic power generation data sequence is less than the K-nearest neighbor distance, and obtain a set of photovoltaic power generation data sequences;

[0082] The maximum value between the K-nearest neighbor distance of the photovoltaic power generation data sequence and the K-nearest neighbor distance of the target photovoltaic power generation data sequence is taken as the reachability distance of the photovoltaic power generation data sequence;

[0083] Based on the reachability distance of the photovoltaic power generation data sequence and the set of photovoltaic power generation data sequences, calculate the local reachability density of the photovoltaic power generation data sequence;

[0084] Specifically, local reachability density measures the average of the inverses of the distances between a photovoltaic (PV) power generation data sequence and other data within its neighborhood. If the distances between the PV power generation data sequence and its neighborhood data are short, its local reachability density will be high; conversely, if these distances are long, the local reachability density of the PV power generation data sequence will be low, indicating that the PV power generation data sequence may be anomalous. The formula for calculating the local reachability density of a PV power generation data sequence is:

[0085] ,

[0086] in, Represents photovoltaic power generation data sequence Locally achievable density; Represents a photovoltaic power generation data sequence The k-neighborhood, i.e., the photovoltaic power generation data sequence Centered on a circle, using photovoltaic power generation data sequences The neighborhood with a radius equal to the K-nearest neighbor distance; Represents photovoltaic power generation data sequence The reachable distance;

[0087] The local outlier factor is calculated based on the local reachability density of the photovoltaic power generation data sequence;

[0088] Specifically, the formula for calculating the local outlier factor of a photovoltaic power generation data sequence is as follows:

[0089] ,

[0090] in, This represents the local outlier factor in a photovoltaic power generation data sequence. Represents photovoltaic power generation data sequence Locally achievable density;

[0091] When the LOF value is close to 1, it means that the density of the photovoltaic power generation data sequence is similar to that of its surrounding neighbors, and it may belong to the same cluster as these neighbors. If the LOF value is less than 1, it means that the photovoltaic power generation data sequence is located in a relatively dense area, and thus may be normal data, and should be given a larger weight. If the LOF value is greater than 1, it usually means that the photovoltaic power generation data sequence is abnormal data, and should be given a smaller weight.

[0092] The weights of the photovoltaic power generation data sequence are obtained by mapping the local outlier factors of the photovoltaic power generation data sequence based on preset coefficients.

[0093] Specifically, the formula for calculating the weights of the photovoltaic power generation data sequence is as follows:

[0094] ,

[0095] in, Indicates the first Weights of individual photovoltaic power generation data sequences; Indicates the preset coefficient;

[0096] By adjusting the values ​​of preset coefficients, local anomaly factors can be mapped to a specific numerical range, thereby allowing for a more accurate measurement of the relative density of the photovoltaic power generation data sequence with its neighboring points, and more effective identification and weighting of anomalous data.

[0097] Furthermore, Gaussian Process Regression (GPR) is a flexible nonparametric statistical method that uses prior information to make state predictions within a Bayesian statistical framework and can estimate the posterior distribution. It can provide not only the mean of the prediction but also the variance and confidence interval, thus enabling the prediction results to express the corresponding uncertainty.

[0098] The Gaussian process regression model constructed in this application, which is weighted by local outlier factors, integrates the local outlier factor algorithm and the Gaussian process regression model. This model analyzes each data point in the historical dataset to assess its degree of anomaly. For data with a high degree of anomaly, the model assigns it a lower weight, thereby reducing the impact of these data on the overall prediction results during the prediction process.

[0099] Specifically, the diagonal weight matrix constructed based on the weights of all photovoltaic power generation data sequences is represented as follows:

[0100] ,

[0101] Furthermore, the weighted Gaussian process is an improvement on Bayesian linear regression. Similar to traditional Gaussian process regression, when constructing a photovoltaic power generation prediction model, the basic form of the model can be set as follows:

[0102]

[0103] in, Represents a linear mapping function; This represents the weight vector; if the model's output contains noise. The weight transformation function can then be expressed as:

[0104] ,

[0105] Combining the weight transformation function and the posterior distribution of the Bayesian linear regression model, the probability distribution of the prediction result obtained by its posterior integral can be expressed as:

[0106] ,

[0107] in, ; Represents the covariance matrix of the prior distribution; The inverse matrix of the covariance matrix of the prior distribution;

[0108] Furthermore, the probability distribution expression of the prediction result is as follows:

[0109] ,

[0110] The joint distribution of the probability distributions of photovoltaic power generation data and forecast data is as follows:

[0111] ,

[0112] Furthermore, the probability distribution model for predicting photovoltaic power generation is expressed as follows:

[0113] ,

[0114] ,

[0115] in, This represents the probability distribution model for predicting photovoltaic power generation. This represents the predicted photovoltaic power generation value output by the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters input into the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters in the photovoltaic power generation data series; This represents the photovoltaic power generation in the photovoltaic power generation data sequence; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; This represents the diagonal weight matrix; Noise indicating photovoltaic power generation output; Represents the identity matrix.

[0116] In Gaussian process regression models, the kernel function directly determines the accuracy of model predictions, and hyperparameters, as an important component of the kernel function, also directly alter the prediction results. In previous Gaussian process regression models, common hyperparameter optimization methods include maximum likelihood estimation (MLE) and the conjugate gradient method. The former is based on the principle of minimizing empirical risk, maximizing a likelihood function composed of samples and unknown parameters, where the probability of the sample occurrence is maximized, thus obtaining the hyperparameters corresponding to the model. The conjugate gradient method introduces conjugacy into the steepest descent method, creating a set of conjugate directions and searching along these directions until the optimal point of the objective function is found. This method accelerates the convergence speed and eliminates the jitter phenomenon of the steepest descent. However, the conjugate gradient method is a deterministic optimization algorithm, which is very sensitive to initial values ​​and is prone to getting trapped in local optima. Therefore, in this embodiment, a genetic algorithm is used to optimize the hyperparameters of the model. The genetic algorithm is a search algorithm that simulates the principles of natural selection and genetics. It is usually used to solve optimization and search problems. It can perform global search in parallel along multiple paths in the solution space, has good global optimization ability, and prevents entering local optima. At the same time, the crossover and mutation operations of the genetic algorithm can be executed in parallel, which improves the efficiency of the algorithm.

[0117] Based on the photovoltaic power generation prediction method provided in the above embodiments, this application also provides a photovoltaic power generation prediction device, such as... Figure 2 As shown, the device specifically includes:

[0118] The data acquisition and correlation coefficient calculation module 10 is used to collect meteorological parameters and photovoltaic power generation at multiple sampling times to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, the correlation coefficient between each meteorological parameter and photovoltaic power generation is calculated.

[0119] Clustering module 20 is used to take K photovoltaic power generation data sequences as cluster centers and calculate the Euclidean distance between each photovoltaic power generation data sequence and each cluster center based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, thereby clustering all photovoltaic power generation data sequences to obtain K initial photovoltaic power generation data clusters.

[0120] The weight calculation module 30 is used to calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and to calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence.

[0121] The model building module 40 is used to construct a diagonal weight matrix based on the weights of all photovoltaic power generation data sequences; and to construct a photovoltaic power generation prediction probability distribution model based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model.

[0122] The parameter optimization and model acquisition module 50 is used to input the multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, optimize the parameters of the photovoltaic power generation prediction probability distribution model, and obtain the target photovoltaic power generation prediction model.

[0123] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the photovoltaic power generation prediction method described above.

[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0128] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for predicting photovoltaic power generation, characterized in that, include: Meteorological parameters and photovoltaic power generation at multiple sampling times were collected to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, the correlation coefficient between each meteorological parameter and photovoltaic power generation was calculated. K photovoltaic power generation data sequences are used as cluster centers. Based on the correlation coefficient between each meteorological parameter and photovoltaic power generation, the Euclidean distance between each photovoltaic power generation data sequence and each cluster center is calculated. Thus, all photovoltaic power generation data sequences are clustered to obtain K initial photovoltaic power generation data clusters. Calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence; Based on the weights of all photovoltaic power generation data sequences, a diagonal weight matrix is ​​constructed; based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model, a probability distribution model for photovoltaic power generation prediction is constructed. The multiple photovoltaic power generation data sequences are input into the photovoltaic power generation prediction probability distribution model, and the parameters of the photovoltaic power generation prediction probability distribution model are optimized to obtain the target photovoltaic power generation prediction model.

2. The photovoltaic power generation prediction method according to claim 1, characterized in that, The formula for calculating the correlation coefficient between each meteorological parameter and photovoltaic power generation is as follows: , in, Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first In the photovoltaic power generation data sequence, the first Various meteorological parameters; This represents the first [number]th ... The average value of various meteorological parameters; Indicates the first Photovoltaic power generation in a photovoltaic power generation data sequence; This represents the average photovoltaic power generation across all photovoltaic power generation data series. Indicates the number of photovoltaic power generation data sequences; , Indicates the types of meteorological parameters.

3. The photovoltaic power generation prediction method according to claim 1, characterized in that, The formula for calculating the Euclidean distance between each photovoltaic power generation data sequence and each cluster center is as follows: , in, Indicates the first A photovoltaic power generation data sequence With the kth cluster center Euclidean distance; Indicates the first The correlation coefficient between various meteorological parameters and photovoltaic power generation; Indicates the first A photovoltaic power generation data sequence The Middle Various meteorological parameters; Represents the k-th cluster center The Middle Various meteorological parameters; , Indicates the types of meteorological parameters.

4. The photovoltaic power generation prediction method according to claim 1, characterized in that, After obtaining K initial photovoltaic power generation data clusters, the following is also included: Based on the class centers of the K initial photovoltaic power generation data clusters, K new cluster centers are obtained; Based on the K new cluster centers, all photovoltaic power generation data sequences are re-clustered to obtain K new photovoltaic power generation data clusters, and it is determined whether the K new photovoltaic power generation data clusters are the same as the K initial photovoltaic power generation data clusters. If the K new photovoltaic power generation data clusters are different from the K initial photovoltaic power generation data clusters, then the class centers of the K new photovoltaic power generation data clusters are used as the K new cluster centers, and all photovoltaic power generation data sequences are re-clustered until the K photovoltaic power generation data clusters obtained by clustering are the same as the K new photovoltaic power generation data clusters.

5. The photovoltaic power generation prediction method according to claim 1, characterized in that, Calculate the local outlier factor for each photovoltaic (PV) power generation data sequence in each initial PV power generation data cluster. Based on the local outlier factor of each PV power generation data sequence, calculate the weight of each PV power generation data sequence, including: For each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, the k-th photovoltaic power generation data sequence that is closest to the photovoltaic power generation data sequence in the initial photovoltaic power generation data cluster is taken as the target photovoltaic power generation data sequence, and the distance between the photovoltaic power generation data sequence and the target photovoltaic power generation data sequence is taken as the K-nearest neighbor distance of the photovoltaic power generation data sequence; Obtain photovoltaic power generation data sequences from the initial photovoltaic power generation data cluster whose distance to the photovoltaic power generation data sequence is less than the K-nearest neighbor distance, and obtain a set of photovoltaic power generation data sequences; The maximum value between the K-nearest neighbor distance of the photovoltaic power generation data sequence and the K-nearest neighbor distance of the target photovoltaic power generation data sequence is taken as the reachability distance of the photovoltaic power generation data sequence; Based on the reachability distance of the photovoltaic power generation data sequence and the set of photovoltaic power generation data sequences, calculate the local reachability density of the photovoltaic power generation data sequence; The local outlier factor of the photovoltaic power generation data sequence is calculated based on the local reachability density of the photovoltaic power generation data sequence. The local outlier factors of the photovoltaic power generation data sequence are mapped based on preset coefficients to obtain the weights of the photovoltaic power generation data sequence.

6. The photovoltaic power generation prediction method according to claim 5, characterized in that, The formula for calculating the local reachability density of a photovoltaic power generation data sequence is: , in, Represents photovoltaic power generation data sequence Locally achievable density; Represents a photovoltaic power generation data sequence The k-neighborhood, i.e., the photovoltaic power generation data sequence Centered on a circle, using photovoltaic power generation data sequences The neighborhood with a radius equal to the K-nearest neighbor distance; Represents photovoltaic power generation data sequence The reachable distance; The formula for calculating the local outlier factor of a photovoltaic power generation data sequence is as follows: , in, This represents the local outlier factor in a photovoltaic power generation data sequence. Represents photovoltaic power generation data sequence Locally achievable density; The formula for calculating the weights of a photovoltaic power generation data sequence is as follows: , in, Indicates the first Weights of individual photovoltaic power generation data sequences; This indicates the preset coefficient.

7. The photovoltaic power generation prediction method according to claim 1, characterized in that, The photovoltaic power generation prediction probability distribution model is expressed as follows: , , in, This represents the probability distribution model for predicting photovoltaic power generation. This represents the predicted photovoltaic power generation value output by the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters input into the photovoltaic power generation prediction probability distribution model; This represents the meteorological parameters in the photovoltaic power generation data series; This represents the photovoltaic power generation in the photovoltaic power generation data sequence; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; express and The covariance matrix; This represents the diagonal weight matrix; Noise indicating photovoltaic power generation output; Represents the identity matrix.

8. The photovoltaic power generation prediction method according to claim 1, characterized in that, After inputting the multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, the parameters of the photovoltaic power generation prediction probability distribution model are optimized using a genetic algorithm.

9. A photovoltaic power generation prediction device, characterized in that, include: The data acquisition and correlation coefficient calculation module is used to collect meteorological parameters and photovoltaic power generation at multiple sampling times to obtain photovoltaic power generation data sequences at multiple sampling times; based on all photovoltaic power generation data sequences, the correlation coefficient between each meteorological parameter and photovoltaic power generation is calculated. The clustering module is used to take K photovoltaic power generation data sequences as cluster centers and calculate the Euclidean distance between each photovoltaic power generation data sequence and each cluster center based on the correlation coefficient between each meteorological parameter and photovoltaic power generation. This allows for the clustering of all photovoltaic power generation data sequences to obtain K initial photovoltaic power generation data clusters. The weight calculation module is used to calculate the local outlier factor of each photovoltaic power generation data sequence in each initial photovoltaic power generation data cluster, and calculate the weight of each photovoltaic power generation data sequence based on the local outlier factor of each photovoltaic power generation data sequence. The model building module is used to construct a diagonal weight matrix based on the weights of all photovoltaic power generation data sequences; and to construct a photovoltaic power generation prediction probability distribution model based on the diagonal weight matrix and the posterior distribution of the Bayesian linear regression model. The parameter optimization and model acquisition module is used to input the multiple photovoltaic power generation data sequences into the photovoltaic power generation prediction probability distribution model, optimize the parameters of the photovoltaic power generation prediction probability distribution model, and obtain the target photovoltaic power generation prediction model.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the photovoltaic power generation prediction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Distributed photovoltaic abnormal data detection method based on periodic optimal path comparison

    CN118551323A

  • Distributed photovoltaic power prediction method, device, equipment, medium and product

    CN119340995A