High-precision short-term photovoltaic power prediction method considering different weather types
By using the Sberman correlation coefficient method and an improved K-means clustering algorithm, combined with Transformer and convolutional neural networks, a high-precision short-term photovoltaic power prediction network was constructed. This solves the problem of insufficient accuracy in photovoltaic power generation prediction in existing technologies, and improves the stability of the power system and the economic benefits of photovoltaic power generation.
Patent Information
- Application Number
- CN202510918522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-28
AI Technical Summary
Existing photovoltaic power generation prediction technologies have shortcomings in feature selection, data classification, and prediction models. They cannot accurately identify nonlinear correlation features and have difficulty distinguishing photovoltaic output power data under different weather conditions, resulting in low prediction accuracy and failing to meet the needs of high-precision practical applications.
The Sperman correlation coefficient method was used to analyze the correlation between meteorological and photovoltaic data, and a stepwise feature subset of meteorological and photovoltaic data was generated. The final dataset was determined by the multi-index coupling correlation entropy screening method, and the improved K-means clustering algorithm was used to achieve accurate clustering of sunny, rainy and cloudy days. A short-term photovoltaic power prediction network integrating Transformer, convolutional neural network and support vector machine was constructed.
It has achieved high-precision short-term photovoltaic power forecasting, improved the reliability of the power system and the stable operation of the power grid, and provided scientific decision support for real-time power dispatching of photovoltaic power plants, optimization of grid load balance and electricity market transactions.
Smart Images

Figure CN120855283A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photovoltaic power generation prediction technology, specifically involving a high-precision short-term photovoltaic power prediction method that takes into account different weather types. It is a precise short-term photovoltaic power generation prediction method based on multi-algorithm fusion, and is particularly suitable for the precise prediction of short-term photovoltaic power generation under different weather types. Background Technology
[0002] With the continuous growth of global demand for clean energy, photovoltaic (PV) power generation, as a sustainable energy solution, is playing an increasingly important role in the energy sector. However, PV power generation is affected by various factors, resulting in significant fluctuations and uncertainties in its output power. This characteristic not only poses significant challenges to the stable operation of power systems, load allocation, and electricity market transactions, but also restricts the efficient integration of PV power generation and the optimal allocation of energy resources. Therefore, achieving high-precision short-term PV power forecasting is crucial for improving power system reliability and ensuring stable grid operation.
[0003] Existing photovoltaic (PV) power generation prediction technologies face numerous unresolved issues. Regarding feature selection, the traditional Pearson correlation coefficient can only effectively identify linear correlation features, failing to accurately filter for the widespread nonlinear correlation features in PV power generation data. This results in an incomplete feature set used for network training, with significant information loss, thus affecting prediction accuracy. In the data classification stage, the traditional K-means clustering algorithm, sensitive to initial cluster centers, easily gets trapped in local optima when classifying PV output power data for different weather types. It cannot accurately distinguish between sunny, rainy, and cloudy weather data, leading to a lack of accuracy and specificity in subsequent network training for different weather types. Regarding prediction network construction, single prediction models have inherent limitations: Support Vector Machines (SVMs) have low computational efficiency and insufficient generalization ability when processing high-dimensional complex data; Convolutional Neural Networks (CNNs) are strong at extracting local features but struggle to capture long-distance dependencies in data; Transformers, while adept at global feature modeling, lack fine-grained characterization of local details. Existing hybrid prediction models often fall short in terms of feature extraction synergy, innovation in fusion mechanisms, and optimization of prediction logic. They are unable to deeply explore the coupling relationship between photovoltaic data features and meteorological conditions, making it difficult to meet the demand for high-precision prediction in practical applications.
[0004] Therefore, there is an urgent need for a method that can accurately and effectively predict short-term photovoltaic (PV) power output to improve power system reliability and ensure stable grid operation. This invention proposes a high-precision short-term PV power prediction method considering different weather types. It utilizes the Sberman correlation coefficient method to analyze the correlation between meteorological data and PV power output data, generating a series of meteorological-PV feature subsets in a stepwise manner. A multi-index coupling correlation entropy screening method is used to determine the final meteorological-PV dataset. An improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernel is used to accurately cluster the final meteorological-PV dataset for three weather types: sunny, rainy, and cloudy. This constructs a short-term PV power prediction network, thereby achieving high-precision short-term PV power prediction considering different weather types. This method is expected to solve the problems existing in the aforementioned prior art, and there are currently no related reports in this area. Summary of the Invention
[0005] The purpose of this invention is to provide a high-precision short-term photovoltaic power prediction method that considers different weather types. This method involves collecting operational datasets related to photovoltaic power, cleaning the data to remove missing and outlier values, performing correlation analysis on the preprocessed dataset using the Sperman correlation coefficient method, and generating a series of meteorological-photovoltaic feature subsets in a stepwise manner. A multi-index coupling correlation entropy screening method is then used to determine the final meteorological-photovoltaic dataset. An improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernel is used to accurately cluster the final meteorological-photovoltaic dataset for three weather types: sunny, rainy, and cloudy. This constructs a short-term photovoltaic power prediction network, thereby achieving high-precision short-term photovoltaic power prediction that considers different weather types, improving power system reliability, and ensuring stable grid operation.
[0006] To achieve the above objectives, the present invention employs the following technical solution: a high-precision short-term photovoltaic power prediction method considering different weather types, characterized by the following specific steps:
[0007] Step S1: Collect the operation dataset related to photovoltaic power: Based on the photovoltaic power plant operation monitoring system and meteorological observation system, collect the operation dataset related to photovoltaic power. The operation dataset includes the original photovoltaic output power data and the original meteorological data. The meteorological data includes a series of data such as irradiance, wind speed, wind direction, temperature, pressure, humidity and actual irradiance.
[0008] Step S2, obtain the preprocessed dataset: clean the operation dataset related to photovoltaic power collected in step S1, remove invalid data with missing values and outliers, and obtain the preprocessed dataset, which includes photovoltaic output power data and meteorological data, with the category of meteorological data remaining unchanged;
[0009] Step S3: Perform correlation analysis using the Sperman correlation coefficient method: Based on the preprocessed dataset from Step S2, calculate the Sperman correlation coefficient between each meteorological data x and the photovoltaic output power data y. The calculation formula is as follows: In the formula, j is the category number of meteorological data in the preprocessed dataset, m is the total number of data samples in the preprocessed dataset, k is the data sample number in the preprocessed dataset, and ρ j x is the correlation coefficient between the j-th category of meteorological data and the photovoltaic output power data. jk y is the value of the j-th category of meteorological data in the k-th data sample of the preprocessed dataset. k It is the photovoltaic output power value of the k-th data sample in the preprocessed dataset, rank(x) jk ) represents the ranking of the j-th category meteorological data in the preprocessed dataset in the k-th data sample, where rank(y) is the ranking of the j-th category meteorological data. k () represents the sorting of photovoltaic output power in the k-th data sample of the preprocessed dataset;
[0010] Step S4: Generate meteorological-photovoltaic feature subsets based on correlation ranking in a stepwise manner: Based on the Sberman correlation coefficient calculated in step S3, sort the meteorological data of each category in descending order according to the absolute value of the correlation coefficient to obtain ordered meteorological data features. From the sorted meteorological data features, select the first meteorological data feature and the corresponding photovoltaic output power to form a feature subset containing one meteorological data feature; select the first two meteorological data features and the corresponding photovoltaic output power to form a feature subset containing two meteorological data features; and so on, until all meteorological data features and the corresponding photovoltaic output power are selected to form a complete feature subset, thus obtaining a series of meteorological-photovoltaic feature subsets containing different numbers of meteorological data features of different categories, all of which are combined with the corresponding photovoltaic output power.
[0011] Step S5: The final meteorological-photovoltaic dataset is determined using the multi-index coupling correlation entropy screening method: For each meteorological-photovoltaic feature subset obtained in step S4, the multi-index coupling correlation entropy E between the photovoltaic output power y and the corresponding meteorological data feature X is calculated, i.e.: In the formula, N is the total number of meteorological data features in each meteorological-photovoltaic feature subset, n is the sequence number of the meteorological data feature in each meteorological-photovoltaic feature subset, and p n It is the probability distribution of the nth meteorological data feature in each meteorological-photovoltaic feature subset, ρ n It is the correlation coefficient between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset, MI(X). n MI(X, y) represents the mutual information between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset.n ,X i ) represents the mutual information between the nth meteorological data feature in each meteorological-photovoltaic feature subset and other meteorological data features, where i is the sequence number of other meteorological data features in each meteorological-photovoltaic feature subset, and VIF n ω1, ω2, and ω3 are the variance inflation factors of the nth meteorological data feature in each meteorological-photovoltaic feature subset, respectively. The correlation entropy E comprehensively reflects the correlation between meteorological data features and photovoltaic power within the feature subset, as well as the information richness of the meteorological data features. Finally, the correlation entropy of each meteorological-photovoltaic feature subset is compared, and the feature subset with the largest correlation entropy is selected as the final meteorological-photovoltaic dataset. The meteorological data features in the final meteorological-photovoltaic dataset are specific meteorological data features.
[0012] Step S6: Based on the improved K-means clustering algorithm using meteorological feature field coupling and dynamic drift kernel, the final meteorological-photovoltaic dataset is accurately clustered for three weather types: sunny, rainy, and cloudy. In the final meteorological-photovoltaic dataset obtained in step S5, let each data point be ξ, and each data point contains the final meteorological data feature vector. And the corresponding final photovoltaic output power value Ф, where B is the number of meteorological data features in the final meteorological-photovoltaic dataset, and each meteorological data feature component These correspond to specific meteorological data characteristics; the final meteorological data feature vector in the final meteorological-photovoltaic dataset. The meteorological characteristic field intensity P is calculated by processing the data. The calculation formula is as follows: In the formula, λ b We assign weights to each meteorological data feature; select initial cluster centers; perform two-dimensional sorting based on meteorological feature field intensity P and final photovoltaic output power value Ф; and select initial cluster centers c corresponding to the three weather types of sunny, rainy, and cloudy days from typical areas of data distribution. o ={c1,c2,c3}; Introducing dynamic drift kernel similarity, its expression is: Where, ||ξ-c o || 2 τ is the Euclidean distance between the data point and the cluster center, γ is the adjustment factor for the Euclidean distance term, exp() represents the power operation of e; after determining the initial cluster centers and the dynamic drift kernel similarity, the iterative clustering process begins. For each data point ξ, based on the dynamic drift kernel similarity S(ξ,c o ), and assign it to the category of the cluster center with the highest similarity to the dynamic drift kernel, that is, ξ=argmax(S(ξ,c o ), o=1,2,3, introduce weighting coefficient η ξThe cluster centers are iteratively updated using the formula: η ξ,o =S(ξ,c o )·(1+α·P), where α is the meteorological characteristic field intensity adjustment coefficient, used to dynamically adjust the influence weight of the meteorological characteristic field intensity P on the update of cluster centers; the updated cluster centers c ' o The calculation formula is: The iteration stops when the cluster center changes the least in two consecutive iterations. The three cluster results are then labeled with weather type to obtain the sunny day dataset, cloudy day dataset, and rainy day dataset.
[0013] Step S7, Constructing a Short-Term Photovoltaic Power Prediction Network: The network uses a parallel input architecture, consisting of a Transformer feature extraction module and a convolutional neural network feature extraction module connected in parallel. The Transformer feature extraction module extracts global features G using a multi-head self-attention mechanism and multiple feedforward layers. The convolutional neural network feature extraction module obtains local features L through convolution operations and pooling layers. A dynamic feature fusion mechanism based on attention-weighted residual fusion is used to process the global features G and local features L to obtain a fused feature vector F. This fused feature vector F is then processed by the global attention matrix A. g and local attention matrix A l We obtain the weighted global feature G by weighting the global feature G and the local feature L respectively. weighted and local features L weighted The calculation formulas are as follows: G weighted =A g V g L weighted =A l V l , where Q g K g V g and Q l K l V l Here, represents the query, key, and value matrices of the global feature G and the local feature L, respectively, where d is the feature dimension. Residual connections and gating mechanisms are introduced to calculate the fused feature vector F, whose calculation formula is: Gate(G,L) = σ(MLP([G;L])), where Gate(G,L) is the gate vector, calculated using a multilayer perceptron (MLP) and the sigmoid function σ, β is the balance coefficient, Conv is the convolution operation, and ⊙ represents element-wise multiplication. This represents the element-wise addition to obtain the fused feature vector F; the fused feature vector F is input into the support vector machine module, mapped to a high-dimensional space to construct a regression model, and thus obtains the short-term photovoltaic power prediction network;
[0014] Step S8: Achieve high-precision short-term photovoltaic power prediction considering different weather types: First, input the sunny day dataset, cloudy day dataset, and rainy day dataset obtained in step S6 into the short-term photovoltaic power prediction network constructed in step S7. Use this network to train on datasets of different weather types. After training is completed, input specific meteorological data features into the trained short-term photovoltaic power prediction network to achieve high-precision short-term photovoltaic power prediction considering different weather types.
[0015] This invention offers the following advantages and benefits: Addressing the issue that the increasing proportion of photovoltaic (PV) power generation in the energy structure, coupled with the significant impact of its intermittent and fluctuating power output on the stable operation of the power system, this invention collects operational datasets related to PV power. After data cleaning to remove missing and outlier values, the preprocessed dataset is analyzed for correlation using the Spelman correlation coefficient method. A series of meteorological-PV feature subsets based on correlation ranking are then generated in a stepwise manner. A multi-index coupling association entropy screening method is employed to determine the final meteorological-PV dataset. Furthermore, an improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernels is used to accurately cluster the final meteorological-PV dataset for three weather types: sunny, rainy, and cloudy, resulting in sunny, cloudy, and rainy datasets. A short-term PV power prediction network integrating Transformer, convolutional neural network, and support vector machine is constructed. The acquired sunny, cloudy, and rainy datasets are input into the constructed short-term PV power prediction network for training. After training, specific meteorological data features are input into the trained short-term PV power prediction network, thereby achieving high-precision short-term PV power prediction considering different weather types. This invention effectively improves data quality and network prediction performance. Its high-precision prediction results can be directly applied to real-time power dispatching of photovoltaic power plants, grid load balance optimization, and electricity market trading strategy formulation, providing scientific and reliable decision support for intelligent management of energy systems and efficient consumption of renewable energy. Attached Figure Description
[0016] Figure 1 This is a flowchart of the overall method of the present invention.
[0017] Figure 2 This is a schematic diagram of the Sberman correlation coefficient results in this invention.
[0018] Figure 3 This is a graph showing the results of the multi-index coupling correlation entropy screening method in this invention.
[0019] Figure 4 This is the precise clustering diagram of the sunny, rainy, and cloudy day datasets in this invention.
[0020] Figure 5 This is a diagram of the short-term photovoltaic power prediction network structure in this invention.
[0021] Figure 6 The graphs show the prediction results of the short-term photovoltaic power prediction network in this invention for sunny (a), cloudy (b), and rainy (c) days.
[0022] Figure 7 This is a graph showing the evaluation index of the prediction results of the short-term photovoltaic power prediction network in this invention for sunny (a), cloudy (b), and rainy (c) days. Detailed Implementation
[0023] The present invention provides a detailed description of a high-precision short-term photovoltaic power prediction method that takes into account different weather types, with reference to the accompanying drawings.
[0024] This invention discloses a high-precision short-term photovoltaic (PV) power prediction method considering different weather types. It collects operational datasets related to PV power, cleans the data to remove missing and outlier values, and uses the Spelman correlation coefficient method to perform correlation analysis on the preprocessed dataset. A series of meteorological-PV feature subsets based on correlation ranking are then generated in a stepwise manner. A multi-index coupling association entropy screening method is used to determine the final meteorological-PV dataset. Finally, an improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernel is used to accurately cluster the final meteorological-PV dataset for three weather types: sunny, rainy, and cloudy, resulting in sunny, cloudy, and rainy datasets. A short-term PV power prediction network integrating Transformer, convolutional neural network, and support vector machine is constructed to provide scientific and reliable decision support for intelligent energy system management and efficient renewable energy consumption.
[0025] like Figure 1 The diagram shows a flowchart of a high-precision short-term photovoltaic power prediction method considering different weather types provided by the present invention. The specific steps are as follows:
[0026] (1) Collect operational datasets related to photovoltaic power: Based on the photovoltaic power plant operation monitoring system and meteorological observation system, collect operational datasets related to photovoltaic power. The datasets include raw photovoltaic output power data and raw meteorological data. The meteorological data include data in categories such as irradiance, wind speed, wind direction, temperature, pressure, humidity and actual irradiance.
[0027] (2) Obtain the preprocessed dataset: Clean the operation dataset related to photovoltaic power collected in step (1) to remove invalid data with missing values and outliers, and obtain the preprocessed dataset. The dataset includes photovoltaic output power data and meteorological data, and the category of meteorological data remains unchanged.
[0028] (3) Correlation analysis using the Sperman correlation coefficient method: Based on the preprocessed dataset in step (2), calculate the Sperman correlation coefficient between each meteorological data x and the photovoltaic output power data y. The calculation formula is as follows: In the formula, j is the category number of meteorological data in the preprocessed dataset, m is the total number of data samples in the preprocessed dataset, k is the data sample number in the preprocessed dataset, and ρ j x is the correlation coefficient between the j-th category of meteorological data and the photovoltaic output power data. jk y is the value of the j-th category of meteorological data in the k-th data sample of the preprocessed dataset. k It is the photovoltaic output power value of the k-th data sample in the preprocessed dataset, rank(x) jk ) represents the ranking of the j-th category meteorological data in the preprocessed dataset in the k-th data sample, where rank(y) is the ranking of the j-th category meteorological data. k () represents the sorting of photovoltaic output power in the k-th data sample of the preprocessed dataset;
[0029] (4) Stepwise generation of meteorological-photovoltaic feature subsets based on correlation ranking: Based on the Sberman correlation coefficient calculated in step (3), the meteorological data of each category are arranged in descending order according to the absolute value of the correlation coefficient to obtain ordered meteorological data features. From the sorted meteorological data features, the first meteorological data feature and the corresponding photovoltaic output power are selected to form a feature subset containing one meteorological data feature; the first two meteorological data features and the corresponding photovoltaic output power are selected to form a feature subset containing two meteorological data features; and so on, until all meteorological data features and the corresponding photovoltaic output power are selected to form a complete feature subset, thereby obtaining a series of meteorological-photovoltaic feature subsets containing different categories of meteorological data features, all of which are combined with the corresponding photovoltaic output power;
[0030] (5) The final meteorological-photovoltaic dataset is determined using the multi-index coupling correlation entropy screening method: For each meteorological-photovoltaic feature subset obtained in step (4), the multi-index coupling correlation entropy E between the photovoltaic output power y and the corresponding meteorological data feature X is calculated, i.e.: In the formula, N is the total number of meteorological data features in each meteorological-photovoltaic feature subset, n is the sequence number of the meteorological data feature in each meteorological-photovoltaic feature subset, and p nIt is the probability distribution of the nth meteorological data feature in each meteorological-photovoltaic feature subset, ρ n It is the correlation coefficient between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset, MI(X). n MI(X, y) represents the mutual information between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset. n ,X i ) represents the mutual information between the nth meteorological data feature in each meteorological-photovoltaic feature subset and other meteorological data features, where i is the sequence number of other meteorological data features in each meteorological-photovoltaic feature subset, and VIF n ω1, ω2, and ω3 are the variance inflation factors of the nth meteorological data feature in each meteorological-photovoltaic feature subset, and ω1, ω2, and ω3 are weighting coefficients. The correlation entropy E comprehensively reflects the correlation between meteorological data features and photovoltaic power within the feature subset, as well as the information richness of the meteorological data features. Finally, the correlation entropy of each meteorological-photovoltaic feature subset is compared, and the feature subset with the largest correlation entropy is selected as the final meteorological-photovoltaic dataset. The meteorological data features in the final meteorological-photovoltaic dataset are specific meteorological data features.
[0031] (6) Based on the improved K-means clustering algorithm of meteorological feature field coupling and dynamic drift kernel, the final meteorological-photovoltaic dataset is accurately clustered for three weather types: sunny, rainy, and cloudy. In the final meteorological-photovoltaic dataset in step (5), let each data point be ξ, and a data point contains the feature vector of the final meteorological data. And the corresponding final photovoltaic output power value Ф, where B is the number of meteorological data features in the final meteorological-photovoltaic dataset, and each meteorological data feature component Each corresponds to a specific meteorological data characteristic.
[0032] Feature vectors of final meteorological data in the final meteorological-photovoltaic dataset The meteorological characteristic field intensity P is calculated by processing the data. The calculation formula is as follows: In the formula, λ b The weights of each meteorological data feature are assigned. Initial cluster centers are selected, and a two-dimensional sort is performed based on the meteorological feature field intensity P and the final photovoltaic output power value Ф. Initial cluster centers c corresponding to sunny, rainy, and cloudy weather types are selected from typical areas of the data distribution. o ={c1,c2,c3}. Introducing a dynamic drift kernel similarity, its expression is: Where, ||ξ-c o || 2τ is the Euclidean distance between the data point and the cluster center, γ is the adjustment factor of the Euclidean distance term, γ is the adjustment factor of the meteorological characteristic field intensity difference term, and exp() represents the power operation of e.
[0033] After determining the initial cluster centers and the dynamic drift kernel similarity, the iterative clustering process begins. For each data point ξ, based on the dynamic drift kernel similarity S(ξ,c... o ), and assign it to the category of the cluster center with the highest similarity to the dynamic drift kernel, that is, ξ=argmax(S(ξ,c o o = 1, 2, 3. Introduce the weighting coefficient η. ξ The cluster centers are iteratively updated using the formula: η ξ,o =S(ξ,c o )·(1+α·P), where α is the meteorological characteristic field intensity adjustment coefficient, used to dynamically adjust the influence weight of the meteorological characteristic field intensity P on the cluster center update. The updated cluster centers c' o The calculation formula is: The iteration stops when the cluster center changes the least in two consecutive iterations. The three cluster results are then labeled with weather type to obtain the sunny day dataset, cloudy day dataset, and rainy day dataset.
[0034] (7) Construction of a short-term photovoltaic power prediction network: The input of this network adopts a parallel input architecture, consisting of a Transformer feature extraction module and a convolutional neural network feature extraction module connected in parallel. The Transformer feature extraction module extracts global features G using a multi-head self-attention mechanism and multiple feedforward layers, while the convolutional neural network feature extraction module obtains local features L through convolution operations and pooling layers. A dynamic feature fusion mechanism based on attention-weighted residual fusion is used to process the global features G and local features L to obtain a fused feature vector F. The global attention matrix A is then used to perform the fusion. g and local attention matrix A l We weight the global feature G and the local feature L separately to obtain the weighted global feature G. weighted and local features L weighted The calculation formulas are as follows: G weighted =A g V g L weighted =A l V l , where Q g K g V g and Q l K l V lThese are the query, key, and value matrices for the global feature G and the local feature L, respectively, where d is the feature dimension. Residual connections and gating mechanisms are introduced to calculate the fused feature vector F, whose calculation formula is: Gate(G,L) = σ(MLP([G;L])), where Gate(G,L) is the gate vector, calculated using a multilayer perceptron (MLP) and the sigmoid function σ, β is the balance coefficient, Conv is the convolution operation, and ⊙ represents element-wise multiplication. This represents element-wise addition to obtain the fused feature vector F. The fused feature vector F is then input into a support vector machine module, mapped to a high-dimensional space, and a regression model is constructed to obtain the short-term photovoltaic power prediction network.
[0035] (8) Achieve high-precision short-term photovoltaic power prediction considering different weather types: First, input the sunny, cloudy, and rainy day datasets obtained in step S6 into the short-term photovoltaic power prediction network constructed in step S7. Then, train the network on datasets of different weather types. After training is completed, input specific meteorological data features into the trained short-term photovoltaic power prediction network to achieve high-precision short-term photovoltaic power prediction considering different weather types.
[0036] like Figure 2 The image shows the results of the Sperman correlation coefficient. The Sperman correlation coefficient method aims to analyze the relationship between meteorological data and photovoltaic output power, and the results are displayed as a heatmap. The degree of correlation between meteorological data and photovoltaic output power data is represented by color bars; the darker the color, the closer the correlation between the meteorological data and photovoltaic output power data.
[0037] Table 1 shows the results of sorting meteorological data by category in descending order of the absolute value of their correlation coefficients. The table shows that the correlation coefficient between actual irradiance and photovoltaic output power is 0.9355, indicating a strong correlation between the two. Furthermore, the correlation coefficients for humidity, temperature, and wind speed are 0.4188, 0.2692, and 0.2278, respectively, showing some correlation. The correlations between pressure and wind direction and photovoltaic output power are weak, with correlation coefficients close to zero.
[0038] Table 1. Correlation coefficients between various types of meteorological data and photovoltaic output power.
[0039]
[0040] like Figure 3The results, shown, are obtained using the multi-indicator coupling correlation entropy screening method. This reflects the relationship between different numbers of meteorological data features and their correlation entropy. The multi-indicator coupling correlation entropy comprehensively reflects the correlation between meteorological data features and photovoltaic power within a feature subset, as well as the information richness of the meteorological data features. The results indicate that when the number of meteorological data features is 5, the feature subset has the highest correlation entropy and is thus selected as the final meteorological-photovoltaic dataset.
[0041] like Figure 4 The image shows the clustering results for sunny, rainy, and cloudy day datasets. It can be seen that the data points for sunny days are concentrated in the region with higher eigenvalues, while the data points for cloudy days are in the middle range of eigenvalues, clearly distinguishing them from the other two weather types. The data points for rainy days are located in the region with lower eigenvalues, with clear boundaries. The distribution areas of the data points for each of the three weather types are clearly defined and relatively independent, indicating that the clustering algorithm can accurately capture the feature differences between different weather types, achieving good clustering results.
[0042] like Figure 5 The diagram shows the structure of a short-term photovoltaic power prediction network. The network employs a parallel input architecture, consisting of a Transformer feature extraction module and a convolutional neural network feature extraction module connected in parallel. The Transformer feature extraction module extracts global features, while the convolutional neural network feature extraction module acquires local features. A dynamic feature fusion mechanism based on attention-weighted residual fusion is used to process the global and local features to obtain a fused feature vector. This fused feature vector F is then input into a support vector machine module to obtain the short-term photovoltaic power prediction network.
[0043] like Figure 6 The image shows the prediction results of the short-term photovoltaic power prediction network for sunny, cloudy, and rainy days. The method of this invention is compared with convolutional neural networks (CNN), support vector machines (SVM), Transformer, Transformer-CNN, and Transformer-SVM. The photovoltaic output power predicted by this invention for sunny, cloudy, and rainy days is closest to the actual values compared to other methods. It can be seen that this invention has high accuracy in high-precision short-term photovoltaic power prediction considering different weather types.
[0044] like Figure 7 The figure shows the evaluation indicators for the short-term photovoltaic power prediction network under sunny, cloudy, and rainy conditions. The mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R²) are calculated. 2This invention is used to quantitatively evaluate its relation to Convolutional Neural Networks (CNN), Support Vector Machines (SVM), Transformer, Transformer-CNN, and Transformer-SVM methods. The mean absolute error (MAE) ranges from [0, +∞], with a smaller value indicating a smaller difference between the predicted and actual values, and higher prediction accuracy. Similarly, the mean absolute error (MAE) also ranges from [0, +∞), reflecting the magnitude of the difference between the predicted and actual values; a smaller MAE indicates that the prediction is closer to the actual value. R 2 The value range is [0,1]. The closer the value is to 1, the better the fit to the data and the better the performance of the prediction network. Results show that this invention exhibits the best performance, with the highest average MSE, MAE, and R among the three datasets for sunny, cloudy, and rainy days. 2 The values were 0.0857, 0.1482 and 0.9821, respectively.
[0045] In summary, this invention provides high-precision short-term photovoltaic power forecasting that takes into account different weather types. It has significant advantages in the field of photovoltaic power forecasting and can be effectively applied to practical engineering scenarios. It is of key significance for improving the reliability of power systems, ensuring the stable operation of power grids, and enhancing the economic benefits of photovoltaic power generation.
[0046] The above embodiments describe the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are only illustrative of the principles of the present invention. Various changes and modifications can be made to the present invention without departing from the scope of the principles of the present invention, and all such changes and modifications fall within the protection scope of the present invention.
Claims
1. A high-precision short-term photovoltaic power prediction method considering different weather types, characterized in that: By collecting operational datasets related to photovoltaic power, missing and outlier values were removed through data cleaning. Correlation analysis was performed on the preprocessed dataset using the Spelman correlation coefficient method, and a series of meteorological-photovoltaic feature subsets were generated in a stepwise manner. The final meteorological-photovoltaic dataset was determined using a multi-index coupling correlation entropy screening method. An improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernels was used to accurately cluster the datasets for three weather types: sunny, rainy, and cloudy, resulting in sunny, cloudy, and rainy datasets. A short-term photovoltaic power prediction network integrating Transformer, convolutional neural network, and support vector machine was constructed. The acquired sunny, cloudy, and rainy datasets were input into the constructed short-term photovoltaic power prediction network for training. After training, specific meteorological data features were input into the trained short-term photovoltaic power prediction network to achieve high-precision short-term photovoltaic power prediction considering different weather types.
2. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 1, characterized in that... The specific process for collecting the operational dataset related to photovoltaic power is as follows: Based on the photovoltaic power plant operation monitoring system and meteorological observation system, the operational dataset related to photovoltaic power is collected. This operational dataset includes the original photovoltaic output power data and the original meteorological data. The meteorological data includes a series of data such as irradiance, wind speed, wind direction, temperature, pressure, humidity and actual irradiance.
3. The high-precision short-term photovoltaic power forecasting method considering different weather types according to claim 2, characterized in that... The specific process for obtaining the preprocessed dataset is as follows: the collected operational dataset related to photovoltaic power is cleaned to remove invalid data with missing or outlier values, resulting in a preprocessed dataset. This dataset includes photovoltaic output power data and meteorological data, with the category of the meteorological data remaining unchanged.
4. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 3, characterized in that... The specific process of correlation analysis using the Sperman correlation coefficient method is as follows: Based on the preprocessed dataset, calculate the Sperman correlation coefficient between each meteorological data x and the photovoltaic output power data y. The calculation formula is as follows: In the formula, j is the category number of meteorological data in the preprocessed dataset, m is the total number of data samples in the preprocessed dataset, k is the data sample number in the preprocessed dataset, and ρ j x is the correlation coefficient between the j-th category of meteorological data and the photovoltaic output power data. jk y is the value of the j-th category of meteorological data in the k-th data sample of the preprocessed dataset. k It is the photovoltaic output power value of the k-th data sample in the preprocessed dataset, rank(x) jk ) represents the ranking of the j-th category meteorological data in the preprocessed dataset in the k-th data sample, where rank(y) is the ranking of the j-th category meteorological data. k ) represents the sorting of photovoltaic output power in the k-th data sample of the preprocessed dataset.
5. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 4, characterized in that... The specific process of generating meteorological-photovoltaic feature subsets based on correlation ranking in a stepwise manner is as follows: Based on the calculated Spelman correlation coefficient, the meteorological data of each category are arranged in descending order according to the absolute value of the correlation coefficient to obtain ordered meteorological data features. From the ranked meteorological data features, the first meteorological data feature and the corresponding photovoltaic output power are selected to form a feature subset containing one meteorological data feature; the first two meteorological data features and the corresponding photovoltaic output power are selected to form a feature subset containing two meteorological data features; and so on, until all meteorological data features and the corresponding photovoltaic output power are selected to form a complete feature subset, thus obtaining a series of meteorological-photovoltaic feature subsets that contain different numbers of meteorological data features of different categories and are all combined with the corresponding photovoltaic output power.
6. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 5, characterized in that... The specific process of determining the final meteorological-photovoltaic dataset using the multi-index coupling correlation entropy screening method is as follows: For each obtained meteorological-photovoltaic feature subset, calculate the multi-index coupling correlation entropy E between the photovoltaic output power y and the corresponding meteorological data features X, i.e.: In the formula, N is the total number of meteorological data features in each meteorological-photovoltaic feature subset, n is the sequence number of the meteorological data feature in each meteorological-photovoltaic feature subset, and p n It is the probability distribution of the nth meteorological data feature in each meteorological-photovoltaic feature subset, ρ n It is the correlation coefficient between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset, MI(X). n MI(X, y) represents the mutual information between the nth meteorological data feature and the photovoltaic output power in each meteorological-photovoltaic feature subset. n ,X i ) represents the mutual information between the nth meteorological data feature in each meteorological-photovoltaic feature subset and other meteorological data features, where i is the sequence number of other meteorological data features in each meteorological-photovoltaic feature subset, and VIF n ω1, ω2, and ω3 are the variance inflation factors of the nth meteorological data feature in each meteorological-photovoltaic feature subset, respectively. The correlation entropy E comprehensively reflects the correlation between meteorological data features and photovoltaic power within the feature subset, as well as the information richness of the meteorological data features. Finally, the correlation entropy of each meteorological-photovoltaic feature subset is compared, and the feature subset with the largest correlation entropy is selected as the final meteorological-photovoltaic dataset. The meteorological data features in the final meteorological-photovoltaic dataset are specific meteorological data features.
7. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 6, characterized in that... The specific process of using the improved K-means clustering algorithm based on meteorological feature field coupling and dynamic drift kernel to accurately cluster the final meteorological-photovoltaic dataset for three weather types (sunny, rainy, and cloudy) is as follows: In the obtained final meteorological-photovoltaic dataset, let each data point be ξ, and each data point contains the final meteorological data feature vector. And the corresponding final photovoltaic output power value Ф, where B is the number of meteorological data features in the final meteorological-photovoltaic dataset, and each meteorological data feature component These correspond to specific meteorological data characteristics; the final meteorological data feature vector in the final meteorological-photovoltaic dataset. The meteorological characteristic field intensity P is calculated by processing the data. The calculation formula is as follows: In the formula, λ b We assign weights to each meteorological data feature; select initial cluster centers; perform two-dimensional sorting based on meteorological feature field intensity P and final photovoltaic output power value Ф; and select initial cluster centers c corresponding to the three weather types of sunny, rainy, and cloudy days from typical areas of data distribution. o ={c1,c2,c3}; Introducing dynamic drift kernel similarity, its expression is: Where, ||ξ-c o || 2 τ is the Euclidean distance between the data point and the cluster center, γ is the adjustment factor for the Euclidean distance term, exp() represents the power operation of e; after determining the initial cluster centers and the dynamic drift kernel similarity, the iterative clustering process begins. For each data point ξ, based on the dynamic drift kernel similarity S(ξ,c o ), and assign it to the category of the cluster center with the highest similarity to the dynamic drift kernel, that is, ξ=argmax(S(ξ,c o ), o=1,2,3, introduce weighting coefficient η ξ The cluster centers are iteratively updated using the formula: η. ξ,o =S(ξ,c o )·(1+α·P), where α is the meteorological characteristic field intensity adjustment coefficient, used to dynamically adjust the influence weight of the meteorological characteristic field intensity P on the update of cluster centers; the updated cluster centers c' o The calculation formula is: The iteration stops when the cluster center changes the least in two consecutive iterations. The three cluster results are then labeled with weather types to obtain the sunny day dataset, cloudy day dataset, and rainy day dataset.
8. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 7, characterized in that... The specific process of constructing a short-term photovoltaic power prediction network is as follows: The network adopts a parallel input architecture, consisting of a Transformer feature extraction module and a convolutional neural network feature extraction module connected in parallel. The Transformer feature extraction module extracts global features G using a multi-head self-attention mechanism and multiple feedforward layers. The convolutional neural network feature extraction module obtains local features L through convolution operations and pooling layers. A dynamic feature fusion mechanism based on attention-weighted residual fusion is used to process the global features G and local features L to obtain a fused feature vector F. This fused feature vector F is then processed by the global attention matrix A. g and local attention matrix A l We obtain the weighted global feature G by weighting the global feature G and the local feature L respectively. weighted and local features L weighted The calculation formulas are as follows: G weighted =A g V g L weighted =A l V l , where Q g K g V g and Q l K l V l Here, G represents the query, key, and value matrices of the global feature G and the local feature L, respectively, and d is the feature dimension. A residual connection and gating mechanism are introduced to calculate the fused feature vector F, whose calculation formula is: Gate(G,L) = σ(MLP([G;L])), where Gate(G,L) is the gate vector, calculated using a multilayer perceptron (MLP) and the sigmoid function σ, β is the balance coefficient, Conv is the convolution operation, and ⊙ represents element-wise multiplication. This represents the element-wise addition to obtain the fused feature vector F; the fused feature vector F is input into the support vector machine module, mapped to a high-dimensional space to construct a regression model, and thus obtains the short-term photovoltaic power prediction network.
9. The high-precision short-term photovoltaic power prediction method considering different weather types according to claim 8, characterized in that... To achieve high-precision short-term photovoltaic power prediction considering different weather types: First, the acquired sunny day dataset, cloudy day dataset, and rainy day dataset are respectively input into the constructed short-term photovoltaic power prediction network. The network is then trained on datasets of different weather types. After training is completed, specific meteorological data features are input into the trained short-term photovoltaic power prediction network to achieve high-precision short-term photovoltaic power prediction considering different weather types.