Hybrid photovoltaic prediction method based on SARIMA-LSTM

By combining SARIMA and LSTM models, the problem of difficulty in accurately modeling seasonal changes in the prior art is solved, and a higher precision photovoltaic power generation prediction is achieved.

CN120163279APending Publication Date: 2025-06-17JIANGSU ANKEREI MICROGRID RES INST CO LTD +2

Patent Information

Application Number
CN202510199476.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing photovoltaic power generation prediction methods are difficult to accurately model seasonal changes, resulting in inaccurate predictions.

Method used

The photovoltaic prediction method based on SARIMA-LSTM mixture is used to model the seasonal data through the SARIMA model, and the time dependence relationship is captured in combination with the LSTM model to make predictions.

Benefits of technology

The accuracy of photovoltaic power generation prediction has been improved, especially in areas with obvious seasonal changes, the daily prediction accuracy can reach more than 95% and the weekly prediction accuracy can reach more than 90%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163279A_ABST
    Figure CN120163279A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid photovoltaic prediction method based on SARIMA-LSTM, and the method comprises the steps: carrying out the calculation according to historical photovoltaic power generation data in an SCADA system, wherein the historical photovoltaic power generation data comprise the power generation power and the accumulated power generation amount, and the meteorological data comprise the irradiance, the temperature, the humidity, the wind speed, the cloud cover and the like, and are obtained from a meteorological station or an open-source data platform; statistics is carried out by combining an SARIMA model with a seasonal prediction function and an LSTM model for processing a time sequence, and a more accurate photovoltaic power generation prediction value is obtained after combined prediction and evaluation. Compared with the prior art, short-term prediction and long-term prediction can be considered, and seasonal change data prediction is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic power generation, and specifically relates to a hybrid photovoltaic prediction method based on SARIMA-LSTM. Background Art

[0002] The power generation of photovoltaic power is greatly affected by weather, day and night, and seasonal changes, and there are intermittent and unstable phenomena, which affect the stability of the power system, leading to problems such as an increase in energy storage demand and an increase in power dispatching difficulty. Whether it is photovoltaic power generation or photovoltaic users, they hope to more accurately predict the photovoltaic power generation situation in the next week or even a month, so as to facilitate the planning of new energy consumption, production scheduling management, peak shaving and valley filling and other measures. Therefore, a photovoltaic power generation prediction method with high accuracy is needed.

[0003] Common photovoltaic prediction methods generally use physical model-based prediction algorithms or statistical methods.

[0004] Physical model-based prediction algorithms generally use irradiance -> power conversion, meteorological models (temperature, humidity, wind speed, cloud cover), and photovoltaic performance models, etc. for power generation prediction. The advantages of such methods are suitable for long-term prediction. The disadvantages are also obvious, with high dependence on meteorological data and high complexity.

[0005] Statistical methods generally construct models based on the statistical laws of historical data and are suitable for short-term prediction. The advantages are simple algorithms, and the disadvantages are limited prediction capabilities for non-linear and sudden changes.

[0006] Chinese Patent with Publication No. CN 112686445 A discloses "A Photovoltaic Power Generation Prediction Method Based on ARIMA-LSTM-DBN", which can perform relatively accurate photovoltaic power generation prediction by combining physical model prediction and statistical methods. However, different regions may have different periodic characteristics, especially seasonal periodic characteristics (collectively referred to as seasonal cycles in this application), and existing prediction methods cannot accurately model different cycles, especially seasonal changes, and the prediction of data with obvious seasonal changes is inaccurate. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects existing in the prior art and provide a hybrid photovoltaic prediction method based on SARIMA-LSTM, which can better model periodicity, especially seasonal changes.

[0008] To achieve the above purpose, the present invention designs a hybrid photovoltaic prediction method based on SARIMA-LSTM, and the prediction method includes: S1. Data preprocessing, including cleaning the original data, filling in missing items, handling abnormal data, and performing standardization; the data includes photovoltaic power generation data and meteorological data. Among them, the photovoltaic power generation data is historical data or real-time monitoring data of a photovoltaic power generation monitoring device (SCADA system), and the meteorological data is historical meteorological data, real-time monitoring data, or meteorological forecast data of the area or region where the photovoltaic power station is located obtained from a meteorological station or an open-source data platform; S2. Construction and prediction of the SARIMA model. The SARIMA model SARIMA(p, d, q)(P, D, Q, s) includes two parts, the non-seasonal part ARIMA(p, d, q) and the seasonal part (P, D, Q, s); where p, d, and q respectively represent the autoregressive order, differencing order, and moving average order of the non-seasonal part, and P, D, Q, and s respectively represent the autoregressive order, differencing order, moving average order, and seasonal length of the seasonal part; the construction and prediction of the SARIMA model include: S21. Parameter initialization to determine the initial SARIMA model; S22. Optimize parameters and verify, including optimizing all parameters of the SARIMA model by means of grid search to determine the preferred model; perform parameter verification on all preferred models. The parameter verification includes checking whether the model residuals are white noise, generally analyzed and judged by calculating the correlation coefficient of the residuals. The set of parameters with the smallest residual correlation is the optimal model parameters for the final prediction; S23. Model training and prediction, using the preprocessed data to train the SARIMA model to generate SARIMA prediction values. The prediction values include the trend part and the non-linear part of the data after the SARIMA model is fitted; S24. Calculate the residuals to generate a prediction residual sequence. The residuals are the difference between the prediction results and the measured results of the training data set; S3. Construction and training of the LSTM model, predicting and evaluating the preprocessed historical data according to time series samples, including: S31. Input data source, perform input-output conversion and data standardization processing according to the time sliding window; the data source includes at least the prediction residuals (linear prediction errors) of the SARIMA model and meteorological data such as irradiance and temperature of the corresponding time series of relevant environmental characteristics; The input-output conversion includes grouping the data source into input values and target values according to the sliding window; S32. Preliminary model construction, including confirming the number of samples, the number of features, confirming the model structure and hyperparameters; S33. Model training, using the SARIMA model prediction values and residuals to train the model; S34. Model Validation and Evaluation. The model is evaluated by calculating the error between the predicted value and the true value, including preliminarily confirming the performance indicators based on at least one of the mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) being lower than the threshold value, and

[0009]

[0010]

[0011] where y i is the actual value of the output variable, is the predicted or estimated value of y i ; S35. Model Optimization. If the model validation and evaluation fails, continuous model optimization is required, including model structure and / or hyperparameter tuning; specific optimization methods include optimizing at least one of the model structure or hyperparameters, and generally using grid search to adjust several hyperparameters, such as the number of LSTM units, learning rate, lag window number, number of LSTM layers, etc.

[0012] Continuously improve the model, train and evaluate until the evaluation criteria are met; S36. Model Output. Use the optimized LSTM model to output the LSTM correction value.

[0013] S4. Combined Prediction and Evaluation. The data is corrected to obtain the final predicted value, and the final predicted value is compared with the actual value, and the residual is calculated again to evaluate the prediction index, so as to optimize the model and correct the prediction parameters in real time; The final predicted value = SARIMA predicted value + LSTM correction value.

[0014] Furthermore, the photovoltaic power generation data at least includes the power generation power and the cumulative power generation, and the meteorological data includes several items such as irradiance, temperature, humidity, wind speed, and cloud cover.

[0015] Furthermore, the standardization includes magnitude normalization and / or mean-variance normalization, and the method for processing abnormal data includes using the moving average method for noise smoothing; For magnitude normalization, the original data is scaled to the interval [0,1]. For any data x, the normalization formula:

[0016] where is the minimum value in the dataset, is the maximum value in the dataset; The mean-variance normalization is to transform the data \(x\) into data with zero mean and unit variance. :

[0017] where and are the mean and variance of the dataset before transformation, respectively. Furthermore, the data preprocessing further includes data splitting (also known as dataset partitioning), that is, the preprocessed data is divided into a training set, a validation set (used for model parameter adjustment and confirming the model architecture), and a test set (used for evaluating the final model performance) according to the time series, and it is ensured that the time series order of each dataset is not disrupted. However, different datasets can have a definite chronological order or there can be cross-overlap of data times.

[0018] Furthermore, in the construction of the SARIMA model, the parameter initialization includes: using the unit root test (ADF, Augmented Dickey-Fuller Test, a common time series analysis tool for testing whether a time series has the unit root feature) to confirm the non-seasonal differencing order \(d\); using seasonal differencing to confirm the seasonal differencing order \(D\); confirming \(q\) and \(Q\) according to the ACF (autocorrelation function). Among them, the ACF of the non-seasonal part cuts off at lag \(q\), so the value range of \(q\) can be judged; the ACF of the seasonal part is obvious at lags \(1s\), \(2s\), or \(3s\), indicating the value range of \(Q\); confirming \(p\) and \(P\) according to the PACF (partial autocorrelation function). Among them, the PACF of the non-seasonal part cuts off at lag \(p\), indicating the value range of \(p\); the PACF of the seasonal part is obvious at lags \(1s\), \(2s\), or \(3s\); confirming \(s\) according to the data frequency. For daily frequency data, \(s = 24\); for monthly frequency data, \(s = 12\).

[0019] Furthermore, the method for confirming the seasonal length \(s\) includes any one of the following methods: ① First, select common physical periods, and the common physical periods include 24h, 7 days, 30 days (or 1 month), 4 quarters, 12 months, or 365 days; ② Find the significant lag points on the period according to the ACF autocorrelation function; ③ Traverse all the values of \(s\) and select the model with the lowest AIC / BIC (Akaike Information Criterion / Bayesian Information Criterion) score, that is, the reasonable value of \(s\).

[0020] Furthermore, the model verification and evaluation also include residuals Analyze and confirm whether relevant data is captured. If the capture is successful, the residual distribution should at least have one of the following characteristics: ① The histogram is close to a normal distribution, ② The mean value of the residuals is close to 0, ③ The correlation of the residuals is white noise, that is, check whether the residuals are randomly distributed. If the residual values fluctuate randomly around 0 without an obvious trend, it is considered that the data is captured. This method is essentially a more rigorous evaluation criterion.

[0021] Furthermore, the construction and training of the LSTM model also include: S36. Model output, combined with real-time data streams (meteorological data and photovoltaic power generation data), to output prediction results.

[0022] Furthermore, the LSTM model captures time-dependent relationships by introducing memory units and gate mechanisms, and corrects the predicted values. The gates include at least one of the input gate, forget gate, and output gate.

[0023] The advantages and beneficial effects of the present invention are as follows: The present invention takes into account both short-term and long-term predictions and is implemented using the statistical model SARIMA and the machine learning model LSTM. The SARIMA model combines the concepts of ARIMA and seasonal differencing, and is more accurate in modeling and predicting seasonal data. At the same time, it combines the machine learning model LSTM (Long Short-Term Memory), which is an improved recurrent neural network (RNN). It memorizes important information and filters out irrelevant information through a gating mechanism, and is specifically used to process time series data and can capture long-term dependencies.

[0024] The combination of the two prediction models can effectively improve the short-term and long-term prediction accuracy, and at the same time solve the problem that the ARIMA model cannot model seasonal change data, improve the prediction accuracy of seasonal data changes, with a daily prediction accuracy of over 95% and a weekly prediction accuracy of over 90%. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of the SARIMA-LSTM hybrid photovoltaic prediction method of the present invention; Figure 2 is a flowchart of the SARIMA model prediction; Figure 3 is a flowchart of the construction and training of the LSTM model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following combines the drawings and embodiments to further describe the specific embodiments of the present invention. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0027] Example 1: A hybrid photovoltaic prediction method based on SARIMA-LSTM is as Figure 1 shown. The prediction method includes: S1. Data preprocessing, including cleaning the original data, filling in missing items, processing abnormal data, and standardizing it; the data includes photovoltaic power generation data and meteorological data. Among them, the photovoltaic power generation data is historical data or real-time monitoring data of a photovoltaic power generation monitoring device (SCADA system), and the meteorological data is historical meteorological data, real-time monitoring data, or meteorological forecast data of the area or region where the photovoltaic power station is located obtained from a meteorological station or an open-source data platform; S2. Construction and prediction of the SARIMA model. The SARIMA model SARIMA(p, d, q)(P, D, Q, s) includes two parts, the non-seasonal part ARIMA(p, d, q) and the seasonal part (P, D, Q, s); where p, d, and q respectively represent the autoregressive order, differencing order, and moving average order of the non-seasonal part, and P, D, Q, and s respectively represent the autoregressive order, differencing order, moving average order, and seasonal length of the seasonal part; as Figure 2 shown. The construction and prediction of the SARIMA model include: S21. Parameter initialization to determine the initial SARIMA model; S22. Optimize parameters and verify, including optimizing all parameters of the SARIMA model by means of grid search to determine the preferred model. After determining the initial parameters, in this example, AIC / BIC (Akaike Information Criterion / Bayesian Information Criterion) is used to compare the AIC / BIC values of each model under different parameter combinations of p, d, q, P, Q, D, and s. The model with an AIC / BIC value less than the AIC / BIC threshold is the preferred model; perform parameter verification on all preferred models. The parameter verification includes checking whether the model residuals are white noise, generally analyzed and judged by calculating the correlation coefficient of the residuals. The set of parameters with the smallest residual correlation is the optimal model parameter for the final prediction; generally, the preprocessed data is used to carry out parameter optimization and verification work. If the residual correlation coefficients of all preferred models are greater than 0.5, it is considered that the optimization fails this time, and the preferred model can be re-determined by adjusting the AIC / BIC threshold or the grid search step size, etc.; S23. Model training and prediction: Use the preprocessed data to train the SARIMA model to generate SARIMA predicted values. The predicted values include the trend data part and the non-linear part after fitting of the SARIMA model. Note that the focus of S22 is to determine the optimal parameter combination and select the optimal model, while S23 is for formal training and prediction. S22 can be called "tuning", and S23 is for "application" to generate predicted data. S24. Calculate the residuals to generate a predicted residual sequence. The residuals are the difference between the predicted results and the measured results of the training data set. Based on the ARIMA model technology, the present invention introduces a seasonal part, which can be used to more accurately predict time series with periodicity and seasonality. For the ARIMA model, please refer to the reference document CN 112686445 A for details.

[0028] S3. Construction and training of the LSTM model: Predict and evaluate the preprocessed historical data according to time series samples, including: S31. Input data source: Perform input-output conversion and data normalization processing according to a time sliding window. The data source includes at least the predicted residuals (linear prediction errors) of the SARIMA model and meteorological data such as irradiance (or radiation amount) and temperature of the corresponding time series and other relevant environmental characteristics. The input-output conversion includes grouping the data source into input values and target values according to a sliding window. For example, in this embodiment, for the residual sequence data, X = [e1, e2, e3, e4, e5, e6, e7, e8]. Take the sliding window length as 4 and the prediction step as 1. For each window, the first 3 values are used as inputs, and the last data is used as the predicted target value, obtaining multiple groups of time series samples: Input sequence X Target value Y [e1, e2, e3] e4 [e2, e3, e4] e5 [e3, e4, e5] e6 [e4, e5, e6] e7 [e5, e6, e7] e8 S32. Preliminary model construction. In this embodiment, in combination with TensorFlow / Keras (Google's open-source machine learning framework), it includes confirming the number of samples and features, confirming the model structure and hyperparameters. The model structure includes the input sliding window length, the number of LSTM cell layers, and the classification of the predicted values output by the fully connected layer. The hyperparameters include the prediction step, the number of LSTM cells, the learning rate, the batch size, the optimizer, etc. Among them, the number of LSTM cells is the number of neurons in each layer, the learning rate is the speed at which the model parameters are updated, the batch size is the number of samples used to calculate the gradient update during each training, and generally, a medium batch size of 32 - 128 is selected.

[0029] In this embodiment, the total number of samples is 10,000, the number of features is 5 (photovoltaic power generation data, temperature, humidity, wind speed, irradiance), the window length is 4, the number of LSTM cell layers is 2, the classification of the predicted values output by the fully connected layer is a regression task (predicting photovoltaic power generation), the prediction step is 1, the number of LSTM cells is 200 units, the learning rate is 0.001, the batch size is 64, and the optimizer selected is Adam (adaptive optimizer); S33. Model training. Use the predicted values and residuals of the SARIMA model for model training. In this embodiment, set up the training model, use the mean squared error (MSE) to confirm the loss function, and the optimizer is Adam; determine the batch size according to the data scale and hardware performance; set up an early stopping mechanism to terminate the training in advance by monitoring the loss of the validation data (such as the validation set) to prevent overfitting; S34. Model validation and evaluation. Evaluate the model by calculating the error between the predicted values and the true values, including preliminarily confirming the performance indicators according to at least one of the mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) being lower than the threshold value, and

[0030]

[0031]

[0032] where, y i is the actual value of the output variable, is the predicted value or estimated value of y i ; In this embodiment, the mean squared error (MSE) is selected as the evaluation index, the threshold value is set to 0.02, and if the MSE is lower than 0.02, it is considered that the model meets the expectations.

[0033] S35. Model Optimization: If the model verification and evaluation fail, continuous model optimization is required, including model structure and / or hyperparameter tuning. Specific optimization methods include optimizing at least one of the model structure or hyperparameters. Generally, the grid search method is used to adjust several hyperparameters, such as the number of LSTM units, learning rate, lag window number, number of LSTM layers, etc. In this embodiment, the number of LSTM layers and the number of units are mainly optimized to improve the prediction accuracy.

[0034] Continuously improve the model, train and evaluate it until the evaluation criteria are met. S36. Model Output: Use the optimized LSTM model to output the LSTM correction value.

[0035] It should be noted that: The residual data after the prediction of SARIMA should theoretically be white noise. However, due to certain reasons, the prediction results of SARIMA may be inaccurate. For example: ① SARIMA is based on a linear hypothesis, and many real data are non-linear; ② There may be more complex hierarchies (combinations of long-term trends / seasonal fluctuations / short-term noises) in the data. Therefore, this invention introduces LSTM to further correct.

[0036] In addition, in practice, since the input data source is a monitoring device (SCADA system), there may be some data mutations at the software or hardware level, resulting in relatively large differences in the data predicted by SARIMA. This invention also combines LSTM to handle such problems.

[0037] That is, LSTM is introduced to solve the influence of non-linear scenarios and fluctuating data in SARIMA on the prediction results.

[0038] S4. Combined Prediction and Evaluation: The data is corrected to obtain the final prediction value, and the final prediction value is compared with the actual value. The residual evaluation prediction index is calculated again to optimize the model and correct the prediction parameters in real time. The key parameters to be corrected in this step of this embodiment include the SARIMA model parameters (p, d, q)(P, D, Q, s) and the time step, number of features, number of LSTM units, number of layers, etc. in the LSTM model. The final predicted value = SARIMA predicted value + LSTM correction value.

[0039] This invention takes into account both short-term and long-term predictions. Based on the statistical model ARIMA, a SARIMA model is established and combined with the machine learning model LSTM, so as to take into account both short-term and long-term terms, especially the seasonal characteristics of photovoltaic data, and achieve more accurate photovoltaic power generation predictions.

[0040] The SARIMA model (Seasonal Autoregressive Integrated Moving Average model) is an optimized model of ARIMA (Autoregressive Integrated Moving Average model), which is specifically used to process time series data with seasonal patterns. The SARIMA model combines the concepts of ARIMA and seasonal differencing, and is more accurate in modeling and predicting seasonal data.

[0041] At the same time, it combines with the machine learning model LSTM (Long Short-Term Memory), an improved Recurrent Neural Network (RNN), which can memorize important information and filter out irrelevant information through a gating mechanism. It is specifically used to process time series data and can capture long-term dependencies.

[0042] The present invention combines the advantages of two prediction models, which can effectively improve the short-term and long-term prediction accuracy. At the same time, it solves the problem that the ARIMA model cannot model seasonal change data and improves the prediction accuracy of seasonal data changes.

[0043] Preferably, the photovoltaic power generation data at least includes power generation power and cumulative power generation, the meteorological data includes several items such as irradiance, temperature, humidity, wind speed and cloud cover, and the sampling interval includes minute, hour or daily level intervals.

[0044] Preferably, the normalization includes magnitude normalization and / or mean-variance normalization, and the method for processing abnormal data includes the method of smoothing noise by using a moving average method; For the magnitude normalization, that is, the original data is scaled to the interval [0,1]. For any data x, the normalization formula:

[0045] where, is the minimum value in the dataset, is the maximum value in the dataset; For the mean-variance normalization, that is, the data x is converted to with zero mean and unit variance :

[0046] where, and are the mean and variance of the dataset before conversion respectively; In this embodiment, mainly the temperature, humidity, wind speed and irradiance data are normalized.

[0047] Preferably, the data preprocessing further includes data splitting (also known as dataset partitioning), that is, the preprocessed data is divided into a training set, a validation set (used for model parameter adjustment to confirm the model architecture), and a test set (used for evaluating the final model performance) according to the time series, and it is ensured that the time series order of each dataset is not disrupted. However, different datasets can have a definite chronological order or there can be an overlap in data time.

[0048] In this embodiment, historical data is partitioned in the ratio of 70% training set, 15% validation set, and 15% test set.

[0049] Preferably, in the construction of the SARIMA model, the parameter initialization includes: The non-seasonal difference order d is confirmed by using the unit root test (ADF, Augmented Dickey-Fuller Test, a common time series analysis tool for testing whether a time series has the characteristics of a unit root). If the p-value (a parameter index in ADF used to determine whether the ADF statistic is significant, automatically calculated by statistical software or a Python library) is greater than 0.05, the sequence is considered a non-stationary sequence and needs to be differenced until the p-value is less than or equal to 0.05. The method for determining the value of d in this embodiment is as follows: if the p-value of the original sequence is less than or equal to 0.05, then d = 0; if the p-value of the original sequence is greater than 0.05, after performing a first-order difference on the original sequence and the p-value is less than or equal to 0.05, then d = 1.

[0050] Seasonal differencing is used to confirm the seasonal difference order D. In this embodiment, after performing seasonal differencing on the preprocessed data sequence, the ADF test method (the same as d above) is used to confirm the seasonal difference order D; Based on the ACF (autocorrelation function), q and Q are confirmed; among them, the ACF of the non-seasonal part cuts off at lag q, so the value range of q can be judged. In this embodiment, generally by observing the decay of the ACF graph, a rapid decline occurs in the first few lags (lags 0, 1, 2), but the decline is not significant at subsequent lags. For example, if the ACF graph shows a significant rapid decline at lags 1 and 2 and reaches 0 at lag 3, then q = 2; The ACF of the seasonal part is obvious at lags 1s, 2s, or 3s, indicating the value range of Q; for the seasonal part, if analyzed according to the monthly distribution, s is taken as 12, that is, there are 12 different trends, then seasonal peaks may appear at 1s (12), 2s (24), 3s (36), etc. on the ACF graph. If there are peaks at 1s and 2s, then Q = 2; Confirm p and P according to the PACF (Partial Autocorrelation Function). For the non-seasonal part, the PACF cuts off at lag p, indicating the range of values for p. Generally, the PACF plot shows peaks in the first few lags (lags 0, 1, 2, 3), and the subsequent lags are close to zero. For example, if the PACF has significant peaks at lags 1 and 2 and is not significant at lag 3, then p = 2. The PACF of the seasonal part is obvious at lags 1s, 2s, or 3s, indicating the possible values of P. For the seasonal part, if s is taken as 12, and there is an obvious peak at 12 and no obvious peak at 24 on the PACF plot, then P = 1. Confirm s according to the data frequency. For daily frequency data, s is 24; for monthly frequency data, s is 12.

[0051] Note that the above are only the initial parameters of the SARIMA model, not the final parameters. The final parameters are obtained during the model training and optimization process, and the selection of the initial values mainly affects the training convergence efficiency.

[0052] Preferably, the method for confirming the seasonal length s includes any one of the following methods: ① First, select common physical periods, which include 24h, 7 days, 30 days (or 1 month), 4 quarters, 12 months, or 365 days. Note that the seasonal length in the present invention does not mean the definite four seasons of the year, but represents a kind of periodicity, and multiple periodic features can be found through training.

[0053] ② Find the significant lag points on the period according to the ACF (Autocorrelation Function). ③ Traverse all the values of s and select the model with the lowest AIC / BIC (Akaike Information Criterion / Bayesian Information Criterion) score, that is, the reasonable value of s.

[0054] Preferably, the model verification and evaluation also include residuals Analyze and confirm whether relevant data is captured. If the capture is successful, the residual distribution should at least have one of the following characteristics: ① The histogram is close to a normal distribution, ② The residual mean is close to 0, ③ The residual correlation is white noise, that is, check whether the residuals are randomly distributed. If the residual values fluctuate randomly around 0 without an obvious trend, it is considered that the data is captured. This method is essentially a more stringent evaluation criterion.

[0055] Preferably, the LSTM model captures time-dependent relationships by introducing memory units and gate mechanisms and corrects the predicted values. The gates include at least one of the input gate, forget gate, and output gate.

[0056] Example 2: The difference from Example 1 is that the construction and training of the LSTM model in this example also include: S36. Model output, combined with real-time data streams (meteorological data and photovoltaic power generation data), outputs the prediction results.

[0057] Example 3: The difference from Example 1 is that in this example, the total number of samples is 8000, the window length is 24 days (increasing the serial port length to better capture long-term dependencies), the number of LSTM cell layers is 3, the output prediction value of the fully connected layer is classified as a continuous output (predicting continuous photovoltaic power generation), the prediction step is 7 days (continuously predicting for 7 days), the mean absolute error (MAE) is selected as the evaluation index, and the threshold value is set to 0.015.

[0058] Example 4: The difference from Example 1 is that in this example, the irradiance, temperature, humidity, wind speed, and cloud cover data in the environmental parameters are simultaneously standardized, and the number of features is 6 (photovoltaic power generation data, irradiance, temperature, humidity, wind speed, and cloud cover).

[0059] The above are only some relatively systematic and comprehensive embodiments of the SARIMA-LSTM hybrid photovoltaic prediction method of the present invention. In fact, for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, such as different proportions of the training set, validation set, and test set, different initial values selected for the parameters in the construction of the SARIMA model, or different methods for selection, other parameter selections of the LSTM and SARIMA models, and so on. These combinations or preferred solutions should also be regarded as the protection scope of the present invention and will not be listed one by one here.

Claims

1. A hybrid photovoltaic prediction method based on SARIMA-LSTM, characterized in that: The prediction method comprises: S1. Data preprocessing, including cleaning the original data, filling in missing items, processing abnormal data, and standardizing; the data includes photovoltaic power generation data and meteorological data; S2. SARIMA model construction and prediction. The SARIMA model SARIMA(p, d, q)(P, D, Q, s) includes a non-seasonal part ARIMA(p, d, q) and a seasonal part (P, D, Q, s); wherein p, d, q represent the autoregressive order, difference order and moving average order of the non-seasonal part, respectively, and P, D, Q, s represent the autoregressive order, difference order, moving average order and seasonal length of the seasonal part, respectively; the SARIMA model construction and prediction include: S21, parameter initialization, determine the initial SARIMA model; S22, optimizing parameters and verifying them, including optimizing all parameters of the SARIMA model by means of grid search, and determining the optimal model; performing parameter verification on all the optimal models, wherein the parameter verification includes checking whether the model residuals are white noise, and a set of parameters with the smallest residual correlation is the optimal model parameters for the final prediction; S23, model training and prediction, using the preprocessed data to train the SARIMA model and generate SARIMA prediction values, wherein the prediction values ​​include the trend data part and the nonlinear part after the SARIMA model is fitted; S24, calculating the residual to generate a prediction residual sequence, wherein the residual is the difference between the prediction result and the measured result of the training data set; S3, LSTM model construction and training, prediction and evaluation of preprocessed historical data according to time series samples, including: S31, input data source, perform input-output conversion and data standardization processing according to the time sliding window; the data source at least includes the prediction residual of the SARIMA model and the irradiance and temperature data of the corresponding time series; The input-output conversion includes distinguishing input values ​​and target values ​​from the data source and grouping them according to sliding windows; S32, preliminary model construction, including confirmation of sample number, feature number, model structure and hyperparameters; S33, model training, using SARIMA model prediction values ​​and residuals for model training; S34, model verification and evaluation, evaluating the model by calculating the error between the predicted value and the true value, including preliminarily confirming the performance index according to at least one of the mean square error (MSE), the root mean square error (RMSE), and the mean absolute error (MAE) being lower than the threshold value; S35, model optimization. If the model validation evaluation fails, it is necessary to continue to optimize the model, including model structure and / or hyperparameter tuning; Continue to improve the model, train and evaluate until it meets the evaluation criteria; S36, model output, using the optimized LSTM model to output the LSTM correction value; S4, combined prediction and evaluation, data correction to obtain the final prediction value, and compare the final prediction value with the actual value, and recalculate the residual evaluation prediction index, so as to optimize the model and correct the prediction parameters in real time; The final predicted value = SARIMA predicted value + LSTM corrected value.

2. A SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1, characterized in that: The photovoltaic power generation data at least includes the power generation and the accumulated power generation, and the meteorological data includes several items of irradiance, temperature, humidity, wind speed and cloud cover.

3. A SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1, characterized in that: The standardization includes value normalization and / or mean variance normalization, and the method for processing abnormal data includes a method for smoothing noise using a sliding average method; The value normalization is to scale the original data to the interval [0,1]. For any data x, the normalization formula is: , in, is the minimum value in the data set, is the maximum value in the data set; The mean variance normalization converts the data x into zero mean and unit variance. : , in, and are the mean and variance of the dataset before transformation.

4. The SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1 is characterized in that: The data preprocessing also includes data segmentation, that is, dividing the preprocessed data into a training set, a validation set and a test set according to the time series.

5. The SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1 is characterized in that: The parameter initialization includes: The unit root test was used to confirm the non-seasonal difference order d; Seasonal differences are used to confirm the seasonal difference order D; According to ACF, q and Q are confirmed; among them, the ACF of the non-seasonal part is cut off at the lag q, so the value range of q can be determined; The ACF of the seasonal part is obvious at lags of 1s, 2s, or 3s, indicating the range of values ​​of Q; Confirm p and P based on PACF, where the PACF of the non-seasonal part cuts off at lag p, indicating the range of p values; The PACF of the seasonal part is obvious at lags of 1s, 2s, or 3s, indicating the range of values ​​of P; Confirm s based on the data frequency. For daily frequency data, s is 24; for monthly frequency data, s is 12.

6. A SARIMA-LSTM hybrid photovoltaic prediction method according to claim 5, characterized in that: The method for confirming the seasonal length s includes any one of the following methods: ① First select a common physical period, which includes 24 hours, 7 days, 30 days or 12 months; ② Find the significant lag point on the cycle based on the ACF autocorrelation function; ③ Traverse all values ​​of s and select the model with the lowest AIC / BIC score, which is the reasonable value of s.

7. The SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1 is characterized in that: The model verification and evaluation also includes residual analysis to confirm whether relevant data is captured. If captured successfully, the residual distribution has at least one of the following characteristics: ① the histogram is close to the normal distribution, ② the residual mean is close to 0, and ③ the residual correlation is white noise.

8. The SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1 is characterized in that: The LSTM model construction and training also includes: S36. Model output, combined with real-time data stream, outputs prediction results.

9. The SARIMA-LSTM hybrid photovoltaic prediction method according to claim 1 is characterized in that: The LSTM model captures time dependencies by introducing memory units and gate mechanisms to correct predicted values, wherein the gates include at least one of an input gate, a forget gate, and an output gate.

Citation Information

Patent Citations

  • Photovoltaic power generation prediction method based on ARIMA-LSTM-DBN

    CN112686445A

Cited By

  • Intelligent energy storage method and system of photovoltaic power station

    CN120978833A

  • Intelligent energy storage method and system for a photovoltaic power plant

    CN120978833B

  • Method for predicting cathode protection potential of FPSO (floating production storage and offloading) riser support structure

    CN121365382A

  • Distributed photovoltaic grid-connected regulation and control system

    CN121863570A