Photovoltaic power prediction method based on RFI and PFE

Through the combination of RFI and PFE, the problem of low accuracy of the photovoltaic power prediction model under small sample conditions is solved. Through feature importance sorting and polynomial up-dimensional optimization, the accuracy and nonlinear relationship capture capability of photovoltaic power prediction are improved.

CN120296427APending Publication Date: 2025-07-11XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437499.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing photovoltaic power prediction methods have low accuracy under small sample conditions and cannot effectively capture nonlinear feature correlations. Moreover, factors such as photovoltaic plate life attenuation and dust accumulation are not reflected in the data set, affecting the prediction accuracy.

Method used

Using a combination of random forest importance sorting (RFI) and polynomial feature engineering (PFE), a final photovoltaic power prediction model is constructed through feature importance sorting and polynomial up-dimensional optimization to improve the nonlinear relationship capture capability of the model.

Benefits of technology

The prediction accuracy of the photovoltaic power prediction model under small sample conditions is improved, and the nonlinear feature changes of the photovoltaic panel can be better captured, which improves the model's prediction ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296427A_ABST
    Figure CN120296427A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic power prediction method based on RFI and PFE, and relates to the technical field of photovoltaic power prediction. Comprising the steps of obtaining a to-be-predicted photovoltaic power data set; inputting the photovoltaic power data set to be predicted into the final power prediction model to obtain a prediction result; the construction method of the model comprises the following steps: determining an input photovoltaic power data set and carrying out data preprocessing to obtain a preprocessed photovoltaic power data set; performing feature importance sorting on the preprocessed photovoltaic power data set, and determining importance measurement to obtain a key feature set; performing polynomial dimension raising optimization on the key feature set to determine a dimension raising feature set; determining an optimal feature number and polynomial dimension raising times according to 30-fold cross validation to obtain an input feature set; determining an initial prediction model and performing hyper-parameter grid optimization; and training the initial prediction model by using the input feature set to obtain a final power prediction model. According to the invention, the problem of low prediction model precision under the small sample condition is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic power generation prediction, and particularly to a photovoltaic power prediction method based on RFI and PFE. Background Art

[0002] With the intensification of traditional energy shortages and environmental pollution problems, photovoltaic power generation, as a clean energy source, has received extensive attention. Among them, photovoltaic power prediction plays a very important role in power grid dispatching and system stability. Existing prediction methods usually rely on a large amount of historical data to train models. However, in actual situations, factors that evolve over time, such as the attenuation of the photovoltaic panel lifespan and changes in surface cleanliness, cannot be directly reflected in the dataset, and the prediction accuracy will thus be limited. At the same time, traditional feature selection methods can effectively capture linear relationships but cannot effectively identify non-linear feature associations, further restricting the model performance under small sample conditions.

[0003] In the machine learning task of photovoltaic power prediction based on numerical weather prediction, the input features only include meteorological factors, photovoltaic panel lifespan, photovoltaic panel dust accumulation, etc., and the negative impact on the photoelectric conversion efficiency becomes larger as time goes by. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a photovoltaic power prediction method based on RFI and PFE, and the present invention solves the problem of low accuracy of the photovoltaic power prediction model under small sample conditions in the prior art.

[0005] To achieve the above purpose, the present invention provides the following solution:

[0006] A photovoltaic power prediction method based on RFI and PFE, comprising:

[0007] Obtain a photovoltaic power dataset to be predicted;

[0008] Input the photovoltaic power dataset to be predicted into the final power prediction model to obtain a prediction result;

[0009] Among them, the construction method of the final power prediction model is:

[0010] Determine the input photovoltaic power dataset and perform data preprocessing to obtain a preprocessed photovoltaic power dataset;

[0011] Perform feature importance ranking on the preprocessed photovoltaic power dataset to determine the importance measure and obtain a key feature set;

[0012] Perform polynomial dimension elevation optimization on the key feature set to determine a dimension elevation feature set;

[0013] Determine the optimal number of features and the degree of polynomial dimensionality increase according to 30-fold cross-validation to obtain the input feature set;

[0014] Determine the initial prediction model and perform hyperparameter grid search optimization, where the initial prediction model is an MLPR model;

[0015] Train the initial prediction model using the input feature set to obtain the final power prediction model.

[0016] Preferably, the determining the input photovoltaic power data set and performing data preprocessing to obtain the preprocessed photovoltaic power data set includes:

[0017] Determine the data set of the time period without missing values in the input photovoltaic power data set to obtain the time-complete data set;

[0018] According to the preset wind speed threshold and the preset wind direction threshold, determine the abnormal data set in the time-complete data;

[0019] Eliminate the abnormal data set to obtain the preprocessed photovoltaic power data set.

[0020] Preferably, performing feature importance ranking on the preprocessed photovoltaic power data set to determine the importance measure to obtain the key feature set includes:

[0021] Construct a random forest regression model, and use the mean squared error as the evaluation index for the decision tree node splitting quality;

[0022] According to the evaluation index, accumulate the difference in MSE before and after each node split in each decision tree, and calculate the importance measure of each feature;

[0023] Determine the importance value according to the importance measure and perform normalization to determine the key feature set.

[0024] Preferably, the calculation expression of the mean squared error is:

[0025]

[0026] where, y i represents the i-th actual value, represents the i-th predicted value, and n is the number of samples.

[0027] Preferably, the performing polynomial dimensionality increase optimization on the key feature set to determine the dimensionality-increased feature set includes:

[0028] Determine the power terms and cross terms;

[0029] Use the power terms and cross terms to perform polynomial dimensionality increase optimization on the key feature set to obtain the dimensionality-increased feature set.

[0030] Preferably, the expression for the number of features in the dimensionality-increased feature set is:

[0031]

[0032] where K is the number of features, m is the number of original features, and n is the polynomial degree.

[0033] The present invention discloses the following technical effects:

[0034] The present invention provides a photovoltaic power prediction method based on RFI and PFE, including: obtaining a photovoltaic power data set to be predicted;

[0035] Inputting the photovoltaic power data set to be predicted into the final power prediction model to obtain a prediction result; wherein, the construction method of the final power prediction model is: determining an input photovoltaic power data set and performing data preprocessing to obtain a preprocessed photovoltaic power data set; performing feature importance ranking on the preprocessed photovoltaic power data set to determine an importance metric to obtain a key feature set; performing polynomial dimensionality-increased optimization on the key feature set to determine a dimensionality-increased feature set; determining an optimal number of features and a polynomial dimensionality-increased degree according to 30-fold cross-validation to obtain an input feature set; determining an initial prediction model and performing hyperparameter grid search, the initial prediction model being an MLPR model; training the initial prediction model using the input feature set to obtain the final power prediction model. The feature ranking method based on RFI in the present invention can capture the non-linear relationship between features and the target better than general feature ranking methods; the dimensionality-increased method based on PFE can mine the non-linear relationship of the original data, thereby improving the prediction ability of the MLPR model which is good at processing high-dimensional data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0037] Figure 1 It is a flowchart of a photovoltaic power prediction method based on RFI and PFE provided by an embodiment of the present invention;

[0038] Figure 2 It is a data mining flowchart provided by an embodiment of the present invention;

[0039] Figure 3 It is a schematic diagram of the principle of cross-validation provided by an embodiment of the present invention;

[0040] Figure 4 Schematic diagram of the generated power curve of the training set provided by the embodiment of the present invention. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0043] As Figure 1 shown, the present invention provides a photovoltaic power prediction method based on RFI and PFE, including:

[0044] Step 100: Obtain the photovoltaic power data set to be predicted;

[0045] Step 200: Input the photovoltaic power data set to be predicted into the final power prediction model to obtain a prediction result;

[0046] Among them, the construction method of the final power prediction model is:

[0047] Step 201: Determine the input photovoltaic power data set and perform data preprocessing to obtain the preprocessed photovoltaic power data set;

[0048] Step 202: Sort the importance of features of the preprocessed photovoltaic power data set to determine the importance measure to obtain the key feature set;

[0049] Step 203: Perform polynomial dimensionality elevation optimization on the key feature set to determine the dimensionality elevation feature set;

[0050] Step 204: Determine the optimal number of features and the polynomial dimensionality elevation times according to 30-fold cross-validation to obtain the input feature set;

[0051] Step 205: Determine the initial prediction model and perform hyperparameter grid search. The initial prediction model is an MLPR model;

[0052] Step 206: Use the input feature set to train the initial prediction model to obtain the final power prediction model.

[0053] Furthermore, as Figures 2 - 3As shown, determining the input photovoltaic power data set and performing data preprocessing to obtain the preprocessed photovoltaic power data set includes:

[0054] Determining the data set of the time period without missing values in the input photovoltaic power data set to obtain a time-complete data set;

[0055] According to the preset wind speed threshold and the preset wind direction threshold, determining the abnormal data set in the time-complete data;

[0056] Eliminating the abnormal data set to obtain the preprocessed photovoltaic power data set.

[0057] Specifically, the data preprocessing step performs outlier detection and processing on the photovoltaic power data set. First, the data set of the time period without missing values is selected, and then the data with a wind speed > 20 m / s and a wind direction < 0 in the training set are identified as abnormal data and replaced with the average value before and after.

[0058] Furthermore, performing feature importance ranking on the preprocessed photovoltaic power data set to determine the importance measure to obtain the key feature set, including:

[0059] Constructing a random forest regression model, with the mean squared error as the evaluation index for the decision tree node splitting quality;

[0060] According to the evaluation index, accumulating the difference in MSE before and after each node split in each decision tree, and calculating the importance measure of each feature;

[0061] Determining the importance value according to the importance measure and performing normalization to determine the key feature set.

[0062] Specifically, the feature importance ranking step first sets the parameters of the random forest regression model. In the present invention, the mean squared error (MSE) is selected to measure the splitting quality in each random decision tree, that is:

[0063]

[0064] In the formula: y i represents the i-th actual value, represents the i-th predicted value.

[0065] Then calculate the difference in MSE before and after each node split of the decision tree, and accumulate the reduced values of MSE before and after the split of all nodes that select the same feature for splitting as the importance measure of this feature.

[0066] Finally, divide the importance value of each feature by the sum of the importance values of all features to ensure that the sum of the importance of all features is 1, and the importance of each feature is between [0, 1].

[0067] Further, the polynomial dimensionality elevation optimization of the key feature set to determine the dimensionality elevation feature set includes:

[0068] Determine the power terms and cross terms;

[0069] Use the power terms and cross terms to perform polynomial dimensionality elevation optimization on the key feature set to obtain the dimensionality elevation feature set.

[0070] Specifically, the polynomial dimensionality elevation optimization step increases the dimension of the feature space by introducing power terms and cross terms, so as to better capture the non-linear relationship in the data and improve the performance of the model. The total number of features K after n - degree polynomial dimensionality elevation of m features is:

[0071]

[0072] For example, after performing 2 - degree polynomial dimensionality elevation on X1 and X2, there are a total of 5 features: X1, X2, X1^ 2 , X2^ 2 , X1X2. The multi - layer perceptron regression machine (MLPR) automatically extracts useful features through multiple non - linear layers for modeling, can learn high - order non - linear features, and thus can handle high - dimensional data sets.

[0073] Furthermore, in the cross - validation step, the 30 - day training set is divided into 30 parts, and the sample size of each part of the data is the same as that of the test set. 30 - fold cross - validation uses each subset as the validation set and the other 29 parts as the training set, and the mean of the 30 results obtained is regarded as the final result.

[0074] The metrics of the model training and testing steps are root mean square error (RMSE), mean absolute error (MAE), and maximum absolute error (MAXE). In the present invention, RMSE is used as the main reference metric. Hyperparameter grid search is performed on MLPR to determine its more suitable parameters as (hidden_layer_sizes=(80,80), alpha = 0.1), and then the RMSE of MLPR after sequentially reducing the number of features using the feature rankings of PCC and RFI is compared to determine the sorting method suitable for each model.

[0075] Specifically, the present invention also provides another embodiment:

[0076] Taking the data set publicly available on the DKASC website as the research object, a small - sample data with a time resolution of 5 minutes from 6:45 to 18:40 every day for a total of 31 days in a month under winter fluctuating weather is selected for research. Among them, the data of the first 30 days is the training set, and the data of the last 1 day is the test set. Referring to Figure 4 , the power generation power fluctuates significantly on 20 days out of the 30 days of sampling.

[0077] After identifying and processing outliers in the data set, different regression models select appropriate numerical preprocessing methods. Then, all features are sorted according to the importance of the random forest. First, without removing any features, the cross-validation results of not performing polynomial dimensionality increase and performing polynomial dimensionality increase 2 to 3 times are recorded. After that, the original features are successively removed in ascending order of correlation, and the previous polynomial dimensionality increase step is repeated until only 1 original feature remains and the polynomial dimensionality increase step is completed, at which point the loop stops. Thus, the cross-validation results have been traversed from all original features to 3 features and from no polynomial dimensionality increase to 3 times of polynomial dimensionality increase. Finally, the number of original features and the polynomial dimensionality increase times when the cross-validation performance is optimal are taken to train and test the regression model.

[0078] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference may be made to each other.

[0079] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A photovoltaic power prediction method based on RFI and PFE, characterized in that, Including: Obtain the photovoltaic power dataset to be predicted; Input the photovoltaic power dataset to be predicted into the final power prediction model to obtain the prediction result; Among them, the construction method of the final power prediction model is: Determine the input photovoltaic power dataset and perform data preprocessing to obtain the preprocessed photovoltaic power dataset; Perform feature importance ranking on the preprocessed photovoltaic power dataset, determine the importance measure to obtain the key feature set; Perform polynomial dimension elevation optimization on the key feature set to determine the dimension-elevated feature set; Determine the optimal number of features and the polynomial dimension elevation times according to 30-fold cross-validation to obtain the input feature set; Determine the initial prediction model and perform hyperparameter grid search. The initial prediction model is the MLPR model; Use the input feature set to train the initial prediction model to obtain the final power prediction model.

2. The photovoltaic power prediction method based on RFI and PFE according to claim 1, characterized in that The determination of the input photovoltaic power dataset and the performance of data preprocessing to obtain the preprocessed photovoltaic power dataset include: Determine the dataset of the time period without missing values in the input photovoltaic power dataset to obtain the time-complete dataset; Determine the abnormal dataset in the time-complete data according to the preset wind speed threshold and the preset wind direction threshold; Eliminate the abnormal dataset to obtain the preprocessed photovoltaic power dataset.

3. A photovoltaic power prediction method based on RFI and PFE according to claim 1, characterized in that, Performing feature importance ranking on the preprocessed photovoltaic power dataset, determining the importance measure to obtain the key feature set, includes: Construct a random forest regression model, and use the mean squared error as the evaluation index for the decision tree node splitting quality; According to the evaluation index, accumulate the MSE difference before and after each node split in each decision tree, and calculate the importance measure of each feature; Determine the importance value according to the importance measure and perform normalization to determine the key feature set.

4. A photovoltaic power prediction method based on RFI and PFE according to claim 3, characterized in that, The calculation expression of the mean squared error is: where y i represents the i-th actual value, represents the i-th predicted value, and n is the number of samples.

5. A photovoltaic power prediction method based on RFI and PFE according to claim 1, characterized in that, The polynomial dimension elevation optimization of the key feature set to determine the dimension-elevated feature set includes: Determine the power terms and cross terms; Use the power terms and cross terms to perform polynomial dimension elevation optimization on the key feature set to obtain the dimension-elevated feature set.

6. The photovoltaic power prediction method based on RFI and PFE according to claim 5, wherein The expression for the number of features in the dimension-elevated feature set is: Where K is the number of features, m is the number of original features, and n is the polynomial degree.