Photovoltaic power prediction method based on improved VMD and Stacking ensemble learning, storage medium and device

By improving the integrated learning method of VMD and Stacking, dynamically adjusting parameters and optimizing the base learner weights, the problem of insufficient adaptability and accuracy of the photovoltaic power prediction model is solved, and higher prediction accuracy and stability are achieved.

CN120262369APending Publication Date: 2025-07-04JIANGSU OCEAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510310100.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing photovoltaic power prediction model is insufficient in different regions. The artificial setting of the punishment factors and modal totals of traditional VMD algorithms leads to improper signal decomposition, poor generalization ability of a single prediction model, and the traditional Stacking method fails to effectively utilize the diversity of the basic learner, resulting in insufficient prediction accuracy and stability.

Method used

By improving the VMD algorithm dynamically adjusting the penalty factor and the total number of modals, combining the kernel principal component analysis method to reduce the dimensionality, and setting up an objective function that automatically adjusts the weight of the basic learner in the Stacking integrated learning model, comprehensively considering diversity and prediction errors, and optimizing the overall performance of the model.

Benefits of technology

It improves the accuracy and stability of photovoltaic power prediction, reduces the risk of overfitting, and enhances the generalization ability and prediction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120262369A_ABST
    Figure CN120262369A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic power prediction method and system based on improved VMD and Stacking ensemble learning, and a storage medium, and the method comprises the following steps: S1, collecting original meteorological data, and carrying out the preprocessing of interpolation and correlation screening, and obtaining preprocessed data; s2, decomposing the preprocessed data by using the improved VMD, and performing dimensionality reduction on the decomposed data by adopting a kernel principal component analysis method to obtain meteorological characteristic data; improving VMD to dynamically update a penalty factor and a mode total number; s3, the meteorological feature data is used for training an improved Stacking integrated learning model, the improved Stacking integrated learning model is provided with an objective function used for autonomously adjusting the weights of a plurality of base learners, and an output result of the Stacking integrated learning model serves as a prediction result; the device can effectively extract the feature information in the data and fully utilize the diversity of the base learner to provide a more accurate prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to power prediction in photovoltaic power generation, and in particular to a photovoltaic power prediction method, storage medium and device based on improved VMD and Stacking ensemble learning. Background Technique

[0002] Energy is the material basis and driving force for the progress of human society and is crucial for a country's economy and security. Photovoltaic power generation is significantly affected by weather conditions, with volatility and intermittency, resulting in greater difficulty in photovoltaic power prediction, bringing new problems to the risk identification and pre-control of distribution network operation, and posing challenges to the stability and security of the power system. Therefore, in order to improve the grid dispatching ability, optimize energy management, and ensure the safe and stable operation of the grid, it is urgent to conduct in-depth research and analysis on the fluctuation characteristics of photovoltaic output and develop a high-precision photovoltaic power prediction method.

[0003] Meteorological data is the basis for accurate photovoltaic power prediction. Different photovoltaic power prediction models also have significant differences in different scenarios. Failing to consider the power generation characteristics of photovoltaics in different regions will affect the accuracy of photovoltaic power generation prediction. Therefore, predicting solar power generation requires analyzing various meteorological variables. By accurately analyzing these meteorological factors, the power output of photovoltaic power generation can be better predicted. The influence degree of each variable on the output is different. If the meteorological factors with weak correlation are not filtered out, the data set may contain a large amount of redundant information, which will have a negative impact on the accuracy of the prediction model. Therefore, it is necessary to deeply mine meteorological information to accurately reflect the actual power generation capacity and ensure the stability and efficient operation of power supply. In addition to the preliminary screening of data, corresponding features in the data need to be extracted for photovoltaic power prediction. In the prior art, the VMD algorithm is used to decompose meteorological data to obtain modal components with strong regularity reflecting the local characteristics of the original signal. However, the penalty factor and the total number of modes are set manually, lacking autonomy and being difficult to adapt to all signals, easily leading to over-decomposition or under-decomposition of signals, and even unable to effectively separate the intrinsic mode functions of signals.

[0004] With the progress of artificial intelligence, machine learning techniques are gradually being used to enhance photovoltaic power generation prediction. Among them, deep learning network architectures have significant advantages. In particular, the LSTM network is very suitable for managing long-term data dependencies, making it very effective in predicting photovoltaic output. However, it is difficult for a single prediction model to adapt to complex and changing prediction scenarios, and an independent LSTM model is prone to problems such as large prediction deviations and poor generalization ability. Therefore, ensemble learning methods have received increasing attention by combining the advantages of various models, and significant progress has been made in related fields. However, traditional stacking methods do not consider the differences in data in the training set and ignore the differences in the performance of different prediction models under different complex conditions, and there are some limitations in improving the model training diversity of the base learners. Therefore, the existing ensemble learning methods are not effective for photovoltaic power prediction either. Summary of the Invention

[0005] Object of the Invention: The object of the present invention is to provide a photovoltaic power prediction method, storage medium and device based on improved VMD and Stacking ensemble learning, which can effectively extract feature information from data and make full use of the diversity of base learners to provide more accurate prediction results.

[0006] Technical Solution: The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to the present invention includes the following steps:

[0007] S1. Collect the original meteorological data and perform preprocessing of interpolation and correlation screening to obtain preprocessed data;

[0008] S2. Decompose the preprocessed data with improved VMD, and use kernel principal component analysis to reduce the dimension of the decomposed data to obtain meteorological feature data; the improved VMD dynamically updates the penalty factor α and the total number of modes K according to the energy and entropy of the data signal.

[0009] S3. Divide the meteorological feature data into multiple training sets and test sets for training the improved Stacking ensemble learning model. The improved stacking ensemble learning model is provided with an objective function for autonomously adjusting the weights of multiple base learners, and the objective function includes the base learner diversity and the prediction error metric function; the output result of the meta-learner in the Stacking ensemble learning model is used as the photovoltaic power prediction result.

[0010] Based on the above technical solution, first preprocess the original meteorological data, improve its resolution through interpolation method, and then eliminate the data with low correlation with photovoltaic output through correlation screening to obtain preprocessed data with high correlation and qualified resolution. Then decompose it through the improved VMD algorithm. The improved VMD algorithm will dynamically update the penalty factor α and the total number of modes K according to the energy and entropy of the signal, so as to avoid the situation of over-decomposition or under-decomposition caused by inappropriate artificial setting of α and K, and even the inability to effectively separate the intrinsic mode functions of the signal. And reducing the dimension of the data after the improved VMD decomposition by using the kernel principal component analysis method can reduce the complexity of the data and avoid data redundancy from reducing the final prediction accuracy; the improved Stacking ensemble learning model for final prediction can autonomously adjust the weights of multiple base learners through the set objective function, and the objective function includes the diversity of base learners and the prediction error index function. Therefore, the objective function can take into account both the diversity of base learners and the prediction error not exceeding the standard. In this way, it can maximize the diversity between base learners, reduce the risk of overfitting, improve the generalization ability, and avoid the prediction error exceeding the standard to ensure the accuracy of the final prediction; therefore, compared with the prior art, the method of the present invention can effectively extract the feature information in the data and make full use of the diversity of base learners to provide more accurate prediction results.

[0011] The computer-readable storage medium storing one or more programs of the present invention includes one or more programs including instructions that, when executed by a computing device, cause the computing device to execute any of the above methods.

[0012] The device of the present invention includes one or more processors, one or more memories, and one or more programs, wherein one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0013] Beneficial effects: Compared with the prior art, the present invention has the remarkable effect that by improving VMD, it can autonomously update the penalty factor and the total number of modes, and can avoid the situations of over-decomposition, under-decomposition or inability to effectively separate the intrinsic mode functions of the signal during the decomposition of data. Reducing the dimension of the data after the improved VMD decomposition by using the kernel principal component analysis method can improve the final prediction accuracy; by setting an objective function including the diversity of base learners and the prediction error index function for the Stacking ensemble learning model, it can reduce the risk of overfitting, improve the generalization ability, and avoid the prediction error exceeding the standard to ensure the accuracy of the final prediction. Description of the Drawings

[0014] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0015] Figure 2 It is a schematic diagram of the training process of the Stacking ensemble learning model;

[0016] Figure 3 It is a schematic diagram of the overall process of the improved Stacking ensemble learning model of the present invention. Specific embodiments

[0017] As shown in the figure, the photovoltaic power prediction method based on improved VMD and Stacking ensemble learning of the present invention includes the following steps:

[0018] S1. Collect the original meteorological data for preprocessing of interpolation and correlation screening to obtain preprocessed data;

[0019] The resolution of photovoltaic data is insufficient, which is likely to affect the system performance and prediction accuracy. Aiming at the problem of too low time resolution of the original meteorological data, the present invention uses a high-precision interpolation method, and its interpolation function is a first-order polynomial. The missing power generation data is calculated based on the existing observation data, and the low-resolution original meteorological data is transformed into high-resolution data, increasing the integrity and continuity of photovoltaic data, reducing the impact of weather fluctuations on the prediction results, and helping to improve the accuracy of the power generation prediction model. The basic principle of the high-precision interpolation method is as follows:

[0020]

[0021] In the formula, x is the interpolation node; y is the output after interpolation, and the subscripts of x and y are the sequence numbers of different coordinates.

[0022] Then, the multi-source original meteorological data after interpolation processing is subjected to Spearman correlation analysis with the photovoltaic output to capture the dependence relationship between the two, exclude the influencing factors with relatively small correlation, help optimize the model and select appropriate features, so as to improve the prediction accuracy. The calculation formula of the Spearman correlation coefficient is

[0023]

[0024] In the formula, x i represents the rank order of the observed value of the i-th meteorological feature, and y i represents the rank order of the observed value of the i-th photovoltaic output; and respectively represent the average rank orders of the meteorological feature and the photovoltaic output; ρ is the Spearman correlation coefficient; N is the number of sampling points.

[0025] S2. Decompose the preprocessed data with improved VMD, and use kernel principal component analysis to reduce the dimension of the decomposed data to obtain meteorological feature data; The improved VMD dynamically updates the penalty factor α and the total number of modes K according to the energy and entropy of the data signal;

[0026] Photovoltaic data contains rich time-varying information. Therefore, effective signal processing techniques are needed to analyze its characteristics, provide more accurate and interpretable frequency-domain decomposition results, and help accurately analyze the operating state of the photovoltaic power generation system. The present invention uses an improved VMD algorithm to perform multi-level decomposition on the preprocessed photovoltaic meteorological time series data, enrich the diversity of input variables, and highlight the local characteristics of the environmental sequence at different time scales.

[0027] The following briefly describes the specific steps of the existing VMD decomposition:

[0028] (1) Perform Hilbert transform on the sub-modalities to obtain the single-sided spectrum s k (t) as:

[0029]

[0030] In the formula, u k (t) is the k-th intrinsic mode function component; δ(t) is the unit impulse function; j is the imaginary unit, and t is the number of signal points.

[0031] (2) Estimate and correct the center frequency of each modal analytic signal according to the obtained spectrum, and shift the spectrum of the modality to the corresponding baseband.

[0032] (3) Estimate the bandwidth from the squared norm of the demodulated signal. In order to minimize the sum of its bandwidth estimates, a variational model with constraint conditions is constructed as follows:

[0033]

[0034] In the formula, K is the total number of modalities; ω k is the center frequency; f(t) is the original signal.

[0035] (4) Introduce a quadratic penalty factor α and a Lagrange multiplier to perform unconstrained processing on the problem, thereby simplifying the solution process. By iteratively updating the variables of each sub-problem, the solution result of the problem is finally obtained.

[0036] Traditional VMD ensures that the frequency range of each modality is relatively compact and independent by optimizing the frequency center, effectively avoiding the problem of modal aliasing. However, the penalty factor α and the total number of modalities K in VMD are manually set. Different signal characteristics have different requirements for parameters, but the default parameters are difficult to adapt to all signals. Improper parameter selection may lead to over-decomposition, under-decomposition of the signal, or even inability to effectively separate the intrinsic mode functions of the signal. Therefore, the present invention uses the energy and entropy value characteristic data of the signal to optimize the parameter selection. The energy of the signal reflects the total amount of signal strength. The energy of the signal is defined as the discrete sum of the square of the signal amplitude, and the specific calculation formula is:

[0037]

[0038] Where E is the energy of the data signal, and x i is the amplitude of the data signal at the i-th point; L is the total number of sampling points of the data signal.

[0039] The entropy value is an index that measures the information complexity or uncertainty of a signal. When the entropy value is high, the signal contains more complex information and greater uncertainty, while when the entropy value is low, the signal is relatively simple and has strong repeatability. The entropy calculation formula of the signal is as follows:

[0040]

[0041] where p i is the probability of the data signal at the i-th point; N sig is the total number of data signals.

[0042] In order to make the adjustment of the penalty factor α and the total number of modes K more scientific, the signal energy and entropy are appropriately normalized, and the parameters are dynamically adjusted using a linear relationship. The normalization formula is as follows:

[0043] E0 = log(E + 1) (7)

[0044]

[0045] where E0 is the normalized energy of the data signal, and H0, H, H max and H min represent the standard entropy, current entropy, maximum entropy, and minimum entropy of the data signal, respectively.

[0046] The dynamic adjustment relationship is as follows:

[0047] α = α0 + c1·E0 + c2·H0 (9)

[0048] K = max(K min , min(K max , round(a·H0 + b·E0))) (10)

[0049] where α0 is the basic penalty coefficient; K max and K min are the upper and lower limits of the number of modes, respectively; c1, c2, a, and b are dynamic parameters adjusted by Bayesian optimization according to minimizing the ratio of the energy and entropy of the data signal; round is a rounding function.

[0050] Signal energy and entropy can reflect the intensity and complexity of a signal. Combining signal features, K and α are dynamically adjusted to make the VMD decomposition result more stable and applicable to non-stationary signals. The improved VMD enhances the adaptability to signals with different features and effectively avoids the problems of over-decomposition or under-decomposition. In addition, more accurate signal features can be extracted, which helps to improve the quality of the input data for the prediction model.

[0051] A large amount of data will increase the complexity of the model and reduce the prediction accuracy. Aiming at the problem of redundancy in multi-source meteorological data, the present invention performs dimensionality reduction on the meteorological feature data after improved VMD decomposition by using KPCA (Kernel Principal Component Analysis) to construct valuable types of feature data. Kernel principal component analysis is a non-linear data processing method that can effectively reduce the dimensionality of data. Assuming the decomposed data is X = {x1, x2, …, x n}, the main steps of KPCA are as follows:

[0052] (1) Preprocess the original data and project it into a high-dimensional space through a mapping function φ to obtain a kernel matrix φ(X) = {φ(x1), φ(x2), …, φ(x n )}.

[0053] (2) Assume that the kernel matrix has been centered, then the covariance matrix is:

[0054]

[0055] In the formula, x n is the meteorological feature data; n is the number of meteorological feature types; C is the covariance matrix.

[0056] (3) Introduce a kernel function K * = φ T φ, perform eigenvalue decomposition on C to generate eigenvalues λ and eigenvectors η.

[0057] (4) According to the magnitudes of the eigenvalues, calculate the contribution rates of each principal component, set a cumulative contribution rate threshold, and in the present invention, the contribution rate is taken as 85%.

[0058] (5) Map the original data onto the selected principal components to obtain the sample data X T .

[0059] X T = η T [φ(x1, λ), φ(x2, λ), …, φ(x n , λ)] T (12)

[0060] In the formula, X T is the meteorological feature data after dimensionality reduction; λ is the eigenvalue; η is the eigenvector.

[0061] Normalize the processed feature data samples to the interval [0, 1] to improve the training efficiency of the model. The normalization calculation formula is

[0062]

[0063] In the formula, P′, P, P max and P min represent the normalized data, the original value, the maximum value, and the minimum value respectively.

[0064] After the processing in step S2 above, meteorological feature data is finally obtained.

[0065] S3. Divide the meteorological feature data into multiple training sets and test sets for training the improved Stacking ensemble learning model. The improved Stacking ensemble learning model has an objective function for autonomously adjusting the weights of multiple base learners. The objective function includes a base learner diversity metric function and a prediction error metric function; use the output result of the meta-learner in the Stacking ensemble learning model as the prediction result.

[0066] Ensemble learning improves the generalization performance and stability by integrating the predictions from multiple base learners. The present invention uses Stacking ensemble learning with a hierarchical structure. By integrating the prediction results of multiple base learners, the overall prediction performance is improved. The Stacking ensemble learning model proposed in the present invention has two layers: the initial layer is multiple models with the meteorological feature data obtained in step S2 as the input, called base learners; the second layer is a model that uses the prediction results of the base learners on the meteorological feature data training set as the training set and the prediction results of the base learners on the meteorological feature data test set as the test set, called the meta-learner. It trains the outputs of all base learners to obtain the final prediction result.

[0067] When constructing a stacked ensemble model, the combination of multiple models will increase the overall complexity of the model. If the meteorological feature data is directly used for training, overfitting may occur. In this case, in order to improve the stability and accuracy of the prediction, cross-validation is usually used to reconstruct the dataset. Specifically: divide the meteorological feature data obtained through step S2 into an initial training set and a test set for model training. Among them, 90% of the meteorological feature data is used as the training set, and 10% is used as the test set. Extract 3 sub-training sets from the divided initial training set. The number of samples in each sub-training dataset is 1 / 3 of the initial training set. That is, divide the initial training set into multiple different small training sets and train multiple base learners to predict the photovoltaic power generation, so as to enhance the diversity of data and model structure and improve the prediction accuracy of photovoltaic power generation.

[0068] The effectiveness of the Stacking model is directly affected by the number of base learners. Usually, 3 - 5 base learners are set in the training of the Stacking ensemble learning model. Therefore, this invention selects Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbor (KNN), and Back Propagation (BP) neural network as base learners, and uses Long Short-Term Memory (LSTM) as the meta-learner to integrate the prediction results of several base learners. LSTM has excellent capabilities in processing sequence data, especially with higher flexibility and accuracy when dealing with complex data. It can integrate the prediction outputs of various types of base learners, combine the advantages of each model, and improve the overall prediction performance.

[0069] Train SVM, RF, KNN, and BP on three sub-training sets respectively, and a total of 12 training results of base learners are obtained. Among them, each base learner is independent. The difference between the prediction result of each base learner and the actual target value on the corresponding sub-training set reflects the fitting degree of the base learner to the data set. Through the comparison of training errors, select the sub-training set with the smallest training error as the optimal sub-training set, and compare the training effects of each base learner on the optimal sub-training set, and select the best three base learners as the best base learners for Stacking ensemble learning.

[0070] By integrating the prediction results of multiple base learners, Stacking can effectively utilize the advantages of different models, achieve higher prediction accuracy than a single model, and make up for the limitations of a single model. However, in the traditional Stacking method, the prediction error is mainly optimized, and the base learners are usually simply assigned the same weight, which may not fully reflect the advantages and disadvantages of each base learner in different stacking scenarios, and ignores the importance of the diversity of base learners for model integration. If the base learners are too similar, the stacking integration effect will be significantly reduced, resulting in information redundancy and reduced generalization ability. Therefore, it is necessary to improve the traditional Stacking method to overcome some inherent limitations to improve the overall performance and robustness of the model.

[0071] Therefore, this invention optimizes the weights and dynamically adjusts the contribution of each base learner to avoid some poor-performance outputs from affecting the final prediction result. By increasing the prediction diversity between base learners, balancing diversity and prediction accuracy, the integrated model becomes more robust and has stronger generalization ability.

[0072] To solve the problem of insufficient model diversity and enhance the prediction accuracy of base learners, an objective function that combines prediction error and base learner diversity is designed to balance the two, and then optimize the overall performance. The weights of multiple base learners are automatically adjusted according to the objective function; the prediction error metric function MSE and the base learner diversity metric function Diversity are expressed as follows:

[0073]

[0074] Wherein, N sa is the number of prediction samples; y i is the predicted value of the i-th sample, is the true value of the i-th sample; Corr(y pred,i , y pred,j ) is the Pearson correlation coefficient between the predicted values of the i-th base learner and the j-th base learner.; MSE is an index of prediction error, i.e., mean square error; Diversity is an index for measuring the diversity among base learners.

[0075] Therefore, considering the prediction error and the diversity of base learners comprehensively, the designed objective function is as follows:

[0076] Loss = τ·MSE + (1 - τ)·Diversity (16)

[0077] Wherein, τ is an adjustable parameter for controlling the weight balance between error and diversity. The value range of τ is (0, 1). To balance the error and diversity, τ takes the value of 0.5 in this embodiment, and other values can also be selected according to actual situations. Prediction error and diversity are two conflicting objectives and it is difficult to optimize them simultaneously. By adjusting τ, a dynamic balance between the two can be achieved.

[0078] The abilities of different base learners to predict different data are unbalanced. Directly averaging and distributing weights may lead to poor results. Through weight optimization, the contribution weights of base learners can be more reasonably distributed. In addition, when the correlation of the prediction results of base learners is too large, the information redundancy increases and the effect of model integration decreases. By maximizing the diversity among base learners, the risk of overfitting can be reduced and the generalization ability can be improved.

[0079] By inputting meteorological feature data into the improved Stacking ensemble learning model constructed by the present invention, the final output result of the model is the photovoltaic power prediction result.

[0080] The present invention improves the prediction accuracy of photovoltaic power generation by enhancing the diversity of data and model structure, uses multiple sub-training sets to enhance the diversity of data, uses a multi-model strategy to select basic models with higher performance, and adaptively distributes weights according to the differences of multiple base learners to optimize the prediction performance and improve the prediction accuracy.

[0081] The computer-readable storage medium storing one or more programs includes one or more programs including instructions that, when executed by a computing device, cause the computing device to execute any one of the methods according to claims 1 to 8.

[0082] The device described in the present invention includes one or more processors, one or more memories, and one or more programs, where the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors. The one or more programs include instructions for performing any of the methods according to claims 1 to 8.

[0083] To verify the performance of the present invention, a dataset from the Global Energy Forecasting Competition 2014 was used for simulation verification. The dataset includes solar energy data and meteorological data from a certain place between 2012 and 2014. It is a 24-hour all-weather dataset with a data resolution of 1 hour to achieve the prediction of the actual photovoltaic power for the next day. A high-precision interpolation method was used to change the original 1-hour resolution data to 15-minute resolution data. Then, the Spearman correlation coefficient was calculated for the photovoltaic output and 12 kinds of meteorological data. Finally, relative humidity, temperature, surface solar radiation, surface thermal radiation, and net solar radiation at the top of the atmosphere were selected as the preprocessed data. By improving the VMD and KPCA methods, the local features of the signal were extracted to obtain meteorological feature data; and the meteorological feature data was divided into 3 sub-training datasets to improve data diversity. Four algorithms, SVM, RF, KNN, and BP, were used as base learners to train on the 3 sub-training sets, and LSTM was used as the meta-learner. The training errors of the 12 base learners on their respective sub-training datasets are shown in Table 1.

[0084] Table 1 Training errors of base learners on different sub-training sets

[0085]

[0086] As can be seen from Table 1, compared with other sub-training datasets, sub-training set 3 has better overall prediction performance for the base learners. Therefore, sub-training set 3 was selected as the optimal sub-training set. By comparing the training effects of each base learner on sub-training set 3, RF, KNN, and BP were finally selected as the best base learners for ensemble learning.

[0087] Sub-training set 3 was selected for comparative experiments. The method of the present invention was compared with traditional LSTM, ELM, and Stacking models. At the same time, to clarify the improvement effect brought by the improvement of VMD and Stacking ensemble learning in the present invention, two additional models, VMD-KPCA-Stacking and I-VMD-KPCA-Stacking, were added. The training set and test set were divided in a ratio of 7:3. To ensure the accuracy of the experimental results, each experiment was carried out 5 times, and finally the prediction results ranked in the median of the prediction accuracy were taken for subsequent analysis and comparison. The photovoltaic power prediction result indicators are summarized in Table 2.

[0088] Summary of Photovoltaic Power Prediction Indicators in Table 2

[0089]

[0090]

[0091] In Table 2, VMD-KPCA-Stacking is a model using traditional VMD, KPCA and Stacking methods, while I-VMD-KPCA-Stacking is a model using improved VMD and traditional KPCA and Stacking methods, and I-VMD-KPCA-I-Stacking is a model using improved VMD, traditional KPCA and improved Stacking methods. It can be seen from Table 2 that the model proposed in the present invention is superior to other models in each evaluation index and has a better prediction effect. Among them, the value of RMSE is 0.014859 MW, which is reduced by 41.2%, 61.3%, 39.6%, 22.7% and 5.5% respectively compared with the other five models; the value of MAPE is 0.0066265 NW, which is reduced by 46.3%, 64.4%, 42.4%, 25.5% and 4.6% respectively compared with the other five models; the value of MAPE is 2.9558%, which is reduced by 4.29%, 7.92%, 4.84%, 2.39% and 1.23% respectively compared with the other five models; the coefficient of determination R2 is also the highest among the four. In summary, compared with other models, the I-VMD-KPCA-I-Stacking model of the present invention has a significant increase in prediction accuracy, a significant decrease in prediction error, and a higher fitting degree between the prediction result and the true value, indicating that the photovoltaic power prediction model proposed in the present invention has a greater improvement in performance compared with other compared prediction models.

Claims

1. A photovoltaic power prediction method based on improved VMD and Stacking ensemble learning, characterized in that It includes the following steps: S1. Collect the original meteorological data for preprocessing of interpolation and correlation screening to obtain preprocessed data; S2. Decompose the preprocessed data using the improved VMD, and perform dimensionality reduction on the decomposed data using the kernel principal component analysis method to obtain meteorological feature data; The improved VMD dynamically updates the penalty factor α and the total number of modes K according to the energy and entropy of the data signal; S3. Divide the meteorological feature data into multiple training sets and test sets for training the improved Stacking ensemble learning model. The improved Stacking ensemble learning model has an objective function for autonomously adjusting the weights of multiple base learners. The objective function includes the base learner diversity and the prediction error metric function; use the output result of the meta-learner in the Stacking ensemble learning model as the photovoltaic power prediction result.

2. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 1, wherein: In step S1, the Spearman correlation coefficient is used to screen the meteorological data with the correlation with photovoltaic output reaching the requirement.

3. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 1, characterized in that: In step S2, the penalty factor α and the total number of modes K in the improved VMD are determined according to the following formulas respectively α = α0 + c1·E0 + c2·H0 K = max(K min , min(K max , round(a·H0 + b·E0))) Among them, α0 is the basic penalty coefficient; K max and K min are respectively the upper and lower limits of the total number of modes; c1, c2, a, and b are dynamic parameters adjusted by Bayesian optimization according to minimizing the ratio of the energy and entropy of the data signal; round is the rounding function; E0 and H0 are respectively the normalized energy and standard entropy of the data signal.

4. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 3, characterized in that: The calculation formula of the said E0 is E0 = log(E + 1) Wherein, E is the energy of the data signal, and the calculation formula is where x i is the amplitude of the data signal at point i; L is the total number of sampling points of the data signal.

5. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 3, characterized in that: The calculation formula of the said H0 is Among them, H, H max and H min respectively represent the current entropy, maximum entropy, and minimum entropy of the data signal; The calculation formula of H is where p i is the probability of the data signal at point i; N sig is the total number of data signals.

6. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 1, characterized in that: In step S3, the objective function of the improved Stacking ensemble learning model for autonomously adjusting the weights of multiple base learners is Loss = τ·MSE + (1 - τ)·Diversity Wherein, τ is an adjustable parameter, MSE is the prediction error metric function of all base learners, and Diversity is the diversity metric function of all base learners; The calculation formula of MSE is Among them, N sa is the total number of predicted samples in all training sets, y i is the predicted value of the i-th sample, is the true value of the i-th sample; The calculation formula of Diversity is Among them, Corr(y pred,i , y pred,j ) is the Pearson correlation coefficient between the predicted values of the i-th base learner and the j-th base learner.

7. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 6, characterized in that: The value range of the said τ is (0, 1).

8. The photovoltaic power prediction method based on improved VMD and Stacking ensemble learning according to claim 1, characterized in that: In step S2, the meteorological data obtained after dimensionality reduction by the kernel principal component analysis method is normalized.

9. A computer-readable storage medium storing one or more programs, characterized in that: It includes one or more programs including instructions, and when the instructions are executed by a computing device, the computing device is caused to execute any one of the methods described in claims 1 to 8.

10. A device, characterized in that: It includes one or more processors, one or more memories, and one or more programs, wherein one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors. The one or more programs include instructions for executing any one of the methods described in claims 1 to 8.

Citation Information

Cited By

  • Power transmission line icing thickness prediction method, system and equipment based on digital-analog dual-drive model, and medium

    CN120508810A