Day-ahead photovoltaic power prediction method, device and system, and storage medium

By optimizing the hyperparameters of the photovoltaic power prediction model using CatBoost model and arithmetic optimization algorithm, combined with similar daily clustering to process historical data, the limitations of traditional photovoltaic power prediction methods in dealing with complex relationships and multivariate inputs are solved, and higher prediction accuracy and model stability are achieved.

CN120073693APending Publication Date: 2025-05-30BEIJING HUISI HUINENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197458.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional PV power prediction methods have limitations when dealing with complex nonlinear relationships and multivariate inputs, making it difficult to capture the rapid changes and seasonal fluctuations of photovoltaic power generation, and are computationally cost-effective and difficult to adapt to different environmental and equipment conditions.

Method used

The CatBoost model is used to optimize the hyperparameters of the photovoltaic power prediction model with arithmetic optimization algorithm (AOA). By acquiring historical photovoltaic power generation data, similar daily clustering is performed, and the trained and optimized model is used to predict the photovoltaic power a few days ago.

Benefits of technology

It improves prediction accuracy and model stability, enhances the ability to capture nonlinear relationships and multivariate characteristics of photovoltaic power generation data, simplifies the parameter tuning process, and improves the stability and reliability of the model under different environments and equipment conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073693A_ABST
    Figure CN120073693A_ABST
Patent Text Reader

Abstract

The invention discloses a day-ahead photovoltaic power prediction method, device and system, and a storage medium. The method comprises the following steps: S1, obtaining historical photovoltaic power generation data; s2, training a CatBoost model according to historical photovoltaic power generation data to obtain a photovoltaic power prediction model; s3, optimizing the hyper-parameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein real-time photovoltaic power generation data are input into the optimized photovoltaic power prediction model for day-ahead photovoltaic power prediction. By adopting the technical scheme of the invention, the limitation of a traditional photovoltaic power prediction method in processing a complex nonlinear relation and multivariable input is solved, and the accuracy and stability of photovoltaic power generation power prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of photovoltaic power generation, and particularly relates to a method and device, a system, and a storage medium for predicting photovoltaic power on a daily basis. Background Art

[0002] With the rapid development of the global economy, the traditional energy system is facing the risk of depletion, and there is an urgent need to initiate a low-carbon emission revolution. As a new power generation mode, photovoltaic power generation has significant advantages such as wide distribution, pollution-free, and inexhaustible, and has become an important means to promote energy transformation and reduce carbon emissions. However, photovoltaic power generation is significantly affected by weather factors such as solar irradiance and temperature, and has disadvantages such as output intermittency, instability, and randomness, which pose a threat to the operation of the power grid. Therefore, a stable and reliable photovoltaic power prediction model is of great significance for ensuring the stable operation of grid connection and local consumption of the power grid.

[0003] Traditional photovoltaic power prediction methods mainly rely on statistical models (such as time series analysis) and physical models (such as physical process simulation based on meteorological data). Although these methods can predict photovoltaic power generation to a certain extent, they have significant limitations in dealing with complex non-linear relationships and multi-variable inputs. For example, time series analysis methods are difficult to capture the rapid changes and seasonal fluctuations of photovoltaic power generation, while physical models rely on accurate meteorological data and complex physical process modeling, with high computational costs and difficulty in adapting to different environmental and equipment conditions. In addition, traditional methods usually ignore the noise and uncertainty in the data, resulting in insufficient accuracy of the prediction results. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and device, a system, and a storage medium for predicting photovoltaic power on a daily basis.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for predicting photovoltaic power on a daily basis, comprising:

[0007] Step S1, obtaining historical photovoltaic power generation data;

[0008] Step S2, training a CatBoost model according to the historical photovoltaic power generation data to obtain a photovoltaic power prediction model;

[0009] Step S3, optimizing the hyperparameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein, the real-time photovoltaic power generation data is input into the optimized photovoltaic power prediction model for predicting photovoltaic power on a daily basis.

[0010] Preferably, the historical photovoltaic power generation data is processed by similarity day clustering into six sub-datasets: winter sunny days, winter cloudy days, winter rainy days, summer sunny days, summer cloudy days, and summer rainy days. Each training dataset includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, and irradiance intensity. The label data includes: historical power generation data.

[0011] Preferably, in step S3, the hyperparameters of the photovoltaic power prediction model are optimized by the AOA algorithm. The hyperparameters include: T, depth, η, λ, num_leaves.

[0012] The present invention also provides a day-ahead photovoltaic power prediction device, including:

[0013] An acquisition module, configured to acquire historical photovoltaic power generation data;

[0014] A training module, configured to train a CatBoost model according to the historical photovoltaic power generation data to obtain a photovoltaic power prediction model;

[0015] An optimization module, configured to optimize the hyperparameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein, the real-time photovoltaic power generation data is input into the optimized photovoltaic power prediction model for day-ahead photovoltaic power prediction.

[0016] Preferably, the acquisition module is further configured to process the historical photovoltaic power generation data by similarity day clustering into six sub-datasets: winter sunny days, winter cloudy days, winter rainy days, summer sunny days, summer cloudy days, and summer rainy days. Each training dataset includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, and irradiance intensity. The label data includes: historical power generation data.

[0017] Preferably, the training module is configured to optimize the hyperparameters of the photovoltaic power prediction model by the AOA algorithm. The hyperparameters include: T, depth, η, λ, num_leaves.

[0018] The present invention also provides a day-ahead photovoltaic power prediction system, including: a memory and a processor. A computer program is stored on the memory and run by the processor. The computer program, when run by the processor, executes the day-ahead photovoltaic power prediction method.

[0019] The present invention also provides a storage medium, on which a computer program is stored. The computer program, when running, executes the day-ahead photovoltaic power prediction method.

[0020] Compared with the prior art, the present invention has the following technical effects:

[0021] 1. Improve prediction accuracy: By combining the Arithmetic Optimization Algorithm (AOA) and the CatBoost model, the hyperparameters of the model are optimized, enhancing the model's ability to capture the non-linear relationships and multi-variable characteristics of photovoltaic power generation data, thereby improving the prediction accuracy.

[0022] 2. Enhance model stability: Utilize the advantages of the CatBoost model in handling categorical features and reducing overfitting, combined with the global optimization ability of AOA, to ensure the stability and reliability of the model under different environmental and device conditions.

[0023] 3. Simplify the parameter tuning process: The global search ability and high convergence speed of the Arithmetic Optimization Algorithm (AOA) can automatically tune the hyperparameters of the CatBoost model, reducing manual intervention and improving the efficiency and accuracy of parameter tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.

[0025] Figure 1 It is the flowchart of the photovoltaic power prediction method for the present invention embodiment;

[0026] Figure 2 It is the clustering result of similar days in winter time;

[0027] Figure 3 It is the clustering result of similar days in summer time;

[0028] Figure 4 It is the CatBoost training process;

[0029] Figure 5 It is the analysis of the determination coefficient of different models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] Example 1:

[0033] As Figure 1 shown, the embodiment of the present invention provides a method for predicting photovoltaic power for the day ahead, including:

[0034] Step 1: Data preprocessing based on similar-day clustering

[0035] (1) Data collection and preliminary cleaning

[0036] In the research on photovoltaic power prediction, it is first necessary to collect and preliminarily clean the data. In the data collection stage, historical power generation data, eight meteorological characteristic data such as solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, and irradiance intensity, as well as relevant time information such as year, month, day, and hour, are obtained from the photovoltaic power station. Then, the data is cleaned. For the missing values and outliers in the data, the mean filling technique is used to fill the missing data, and the outliers are identified and corrected or deleted through Z-score to ensure the accuracy and integrity of the data set.

[0037] (2) Data standardization

[0038] After the preliminary data cleaning, data standardization processing is carried out. For meteorological characteristics covering irradiance, temperature, humidity, etc., for data with different orders of magnitude and units, it is necessary to convert the characteristics to the same scale through min-max normalization to avoid some characteristics having too great an impact on the model due to too large a numerical range.

[0039]

[0040] Among them, x is the feature column in the data set, and x' is the value after standardization; min(x) represents the minimum value of the feature, and max(x) represents the maximum value of the feature.

[0041] (3) Similar-day clustering

[0042] To achieve refined processing of photovoltaic power generation data, the original dataset is segmented based on the method of similar-day clustering. The K-means clustering algorithm is selected to cluster the days in historical data according to feature similarity, and the dates with similar meteorological conditions and power generation patterns are identified. First, the original dataset is divided into winter time (October - March) and summer time (April - September) according to time information, and then it is divided into three scenarios: sunny days, cloudy days, and rainy / snowy days based on relative humidity and cloud type using the K-means clustering algorithm. Generally speaking, the training dataset is divided into six sub-datasets through the clustering algorithm. As shown in Figure 2 for winter time, it includes: sunny days in winter time, cloudy days in winter time, rainy days in winter time. The clustering results for summer time are as shown in Figure 3 , including sunny days in summer time, cloudy days in summer time, rainy days in summer time. Each training dataset includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, irradiance intensity. The label data: historical power generation data. This step not only helps to understand the similarity between different days, but also enables more targeted model training and improves the model prediction speed.

[0043] Step 2: Construct a Catboost photovoltaic prediction model

[0044] After dividing the training dataset, input the CatBoost algorithm to train the CatBoost photovoltaic power prediction model. Its main idea is to linearly combine weak classifiers, which is an improvement on the Gradient Boosting Decision Tree (GBDT) algorithm. The CatBoost training process is as shown in Figure 4 .

[0045] 1. Processing of categorical features

[0046] For categorical features, the target statistics method is used for encoding. The so-called categorical features, such as cloud type, the specific values of this feature are fixed labels, for example, 0 represents clear sky, 1 represents almost clear sky, 2 represents partly cloudy, 3 represents cloudy. Since this type of categorical label data is used in the calculation process, when the feature distribution of an individual case or the test set is inconsistent with that of the training set, it will lead to the problem of label leakage, that is, the value of x i,k is equal to the label value. To avoid the above problems, the CatBoost algorithm proposes the target statistics encoding method. Target encoding replaces the category according to the average label value of each sample in the entire dataset to obtain the corresponding data feature x i,k , as shown in the following formula:

[0047]

[0048] where: i and j are the row numbers; k is the column number; [·] is the Iverson bracket; y j is the label value.

[0049] Compared with the commonly used one-hot encoding, when facing high-cardinality categories, the corresponding feature dimensions and the depth of the tree will also be significantly reduced, reducing the time and computing resources consumed by the model.

[0050] 2. Model Initialization

[0051] Initialize the model. The predicted value of the model is a constant, usually the average of the target values in the training set. Then set the number of iterations T. CatBoost constructs multiple trees through multiple rounds of iteration, and a new decision tree is constructed in each round of iteration.

[0052] 3. Sorting Boosting Mechanism

[0053] The sorting boost of the training data is an important technique for reducing gradient bias and dealing with feature skewness problems. The sorting boost simulates the process of online learning by randomly permuting the order of the training data, thereby improving the robustness and generalization ability of the model. The following is a detailed explanation and steps:

[0054] ① Construct the split points of the current tree

[0055] Construct the split points of the current tree. For each node n, select the best feature and split point, calculate the gradient of the loss function, and the calculation method is shown in the following formula. Select the split point that minimizes the loss function.

[0056] The definition of the loss function is:

[0057]

[0058] where: is the predicted PV output value of the model, and y is the label value (i.e., historical PV data).

[0059] The gradient calculation formula is:

[0060]

[0061] ② Calculate the values of the leaf nodes

[0062] Then, calculate the values of the leaf nodes. Assuming that the structure of the tree is determined, each leaf node j contains a set of samples {i∣q(x i ) = j}, where q(x i ) is the leaf node to which the sample x i is assigned. The total gradient on the child node is:

[0063]

[0064] Among them, is the gradient of sample i.

[0065] Calculate the total second-order derivative on the leaf node:

[0066]

[0067] Among them, is the second-order derivative of sample i.

[0068] ③ Update the model prediction value

[0069] First, update the prediction value of the leaf node:

[0070]

[0071] Among them, λ is the regularization parameter, which is used to control the smoothness of the leaf node prediction value and avoid overfitting.

[0072] After each round of iteration is completed, add the prediction value of the current tree to the prediction value of the existing model to jointly calculate the new model prediction value Let the prediction value of the current tree be f t (x), then the update formula is:

[0073]

[0074] Among them, η is the learning rate, which is used to control the step size of each update.

[0075] Finally, check whether the iteration stop condition is satisfied. If the preset number of iterations T is reached, exit the loop and complete the training of the CatBoost photovoltaic power prediction model.

[0076] Step 3: Optimize the prediction model with AOA

[0077] Hyperparameter tuning is a key step in optimizing the performance of machine learning models. Appropriate selection of hyperparameters can significantly improve the efficiency and performance of the model. Reasonable hyperparameters can enable the model to converge faster, avoid overfitting or underfitting, and improve the generalization of the model to unseen data.

[0078] The Arithmetic Optimization Algorithm (AOA) is a population-based metaheuristic algorithm. Inspired by the hierarchical structure of arithmetic operators and their dominance from outside to inside, it determines the best element that meets specific criteria from a set of candidate alternatives. As shown in Table 1 below, the CatBoost model has five hyperparameters that need to be optimized. The AOA algorithm can solve the optimization problem without calculating derivatives, and its algorithm process is as follows.

[0079] Table 1

[0080] Hyperparameter Name Chinese Name Optimization Range T Number of Iterations [50,500] depth Depth of the Tree [4,10] η Learning Rate [0.02,0.2] λ L2 Regularization Coefficient [0,10] num_leaves Maximum Number of Leaf Nodes [0,50]

[0081] The optimization process of AOA consists of three stages: initialization, exploration, and exploitation. In the initialization stage, the AOA parameters and the initial random solution are mainly set. In the exploration stage, the multiplication operator or division operator with large-order variation is used to explore the entire space, and a large step size is used to find a more superior position. In the exploitation stage, the addition operator or subtraction operator that is easier to approach the target is used for local optimization and to maintain the diversity of candidate solutions. In addition, excellent performance of the algorithm requires an appropriate balance between exploration and exploitation.

[0082] ①Initialization

[0083] First, AOAQ generates a uniformly distributed random population. The initial population can be obtained by the following formula

[0084] X(i,j) = Rand × (Ub - Lb) + Lb (8)

[0085] where Ub is the upper bound, Lb is the lower bound, Rand is a random number between [0,1], and X(i,j) is the position of the i-th solution in the j-th dimensional space

[0086] ②Mathematical function accelerator MOA

[0087] Judge whether to perform global exploration or local exploitation according to the value of MOA:

[0088]

[0089] where, r 1 is a random number between [0,1], MOA(t) is the current value of the acceleration function, Min is the minimum value of the acceleration function, Max is the maximum value of the acceleration function, T is the maximum number of iterations, and t is the current number of iterations; when r

[0090] < MOA(t), the function will enter the global exploration stage, otherwise it will enter the local exploitation stage. 1

[0091] ③Exploration stage

[0092] Decide whether to adopt the multiplication or division strategy according to the random number r 2 The two strategies have a high degree of dispersion, which is conducive to the particles exploring in the algorithm space. The calculation formula is as follows:

[0093]

[0094] where, r 2 is a random number between [0,1], X(t + 1) is the position of the next-generation particle, X b(t) is the position of the current best fitness particle, μ is the search process control coefficient (with a value of 0.499), ε is the minimum value, and MOP is the mathematical optimizer probability. Its calculation formula is as follows:

[0095]

[0096] Among them, MOP(t) is the current mathematical optimizer probability, α is the iteration sensitivity coefficient, and the higher the value of α, the greater the impact of the number of iterations on MOP(t).

[0097] ④ Development stage

[0098] The arithmetic optimization algorithm conducts local development through addition strategies and subtraction strategies. Both strategies have significant low dispersion and can easily approach the target, which is beneficial for the algorithm to find the

[0099] optimal solution more quickly. Its calculation formula is as follows:

[0100]

[0101] Among them, r 2 is a random number in [0, 1].

[0102] Step 4 Evaluation of the accuracy of photovoltaic power prediction

[0103] This study uses the optimal goodness of fit (R 2 ), also known as the coefficient of determination, to measure the prediction accuracy. The value range of R 2 is from 0 to 1. The closer the value of R 2 is to 1, the closer the predicted value is to the observed value; conversely, the smaller the value of R 2 , the worse the fitting degree of the predicted value to the observed value. The calculation formula is shown in Equation (2-11):

[0104]

[0105] Case study: This paper uses the photovoltaic output and historical meteorological data of a certain place for one year, with a time scale of 1h. Based on clustering according to seasons and weather conditions, 80% is the training set and 20% is the test set. To verify the superiority of the proposed algorithm in photovoltaic day-ahead power prediction, multi-model prediction and comparative analysis are carried out on data in different seasons. In addition to the CatBoost algorithm proposed in this patent, 7 other common models are selected for accuracy comparison: FxtraTrees, XGB, Voting, GBR, DT, SVR, AdaBoost.

[0106] As Figure 5 shown, under different weather conditions, the coefficient of determination of the CatBoost model shows significant superiority. Specifically, the R under winter sunny and summer sunny conditions2 They are 0.962 and 0.944 respectively. Even in complex weather conditions such as rainy days in winter and cloudy days in summer, CatBoost still maintains a high R 2 , which are 0.853 and 0.884 respectively, superior to other traditional models.

[0107] Example 2:

[0108] The embodiment of the present invention also provides a day-ahead photovoltaic power prediction device, including:

[0109] An acquisition module, configured to acquire historical photovoltaic power generation data;

[0110] A training module, configured to train a CatBoost model based on the historical photovoltaic power generation data to obtain a photovoltaic power prediction model;

[0111] An optimization module, configured to optimize the hyperparameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein, the real-time photovoltaic power generation data is input into the optimized photovoltaic power prediction model for day-ahead photovoltaic power prediction.

[0112] As an implementation manner of the embodiment of the present invention, the acquisition module is further configured to perform similarity-day clustering processing on the historical photovoltaic power generation data and divide it into six sub-datasets: sunny days in winter time, cloudy days in winter time, rainy days in winter time, sunny days in summer time, cloudy days in summer time, rainy days in summer time. Each training data set includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, irradiance intensity. The label data includes: historical power generation data.

[0113] As an implementation manner of the embodiment of the present invention, the training module is configured to optimize the hyperparameters of the photovoltaic power prediction model through the AOA algorithm. The hyperparameters include: T, depth, η, λ, num_leaves.

[0114] Example 3:

[0115] The embodiment of the present invention also provides a day-ahead photovoltaic power prediction system, including: a memory and a processor. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, it executes the day-ahead photovoltaic power prediction method.

[0116] Example 4:

[0117] The embodiment of the present invention also provides a storage medium. A computer program is stored on the storage medium. When the computer program runs, it executes the day-ahead photovoltaic power prediction method.

[0118] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for predicting photovoltaic power on the day ahead, characterized in that: include: Step S1, obtaining historical photovoltaic power generation data; Step S2: training a CatBoost model based on historical photovoltaic power generation data to obtain a photovoltaic power prediction model; Step S3, optimizing the hyperparameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein, real-time photovoltaic power generation data is input into the optimized photovoltaic power prediction model to perform day-ahead photovoltaic power prediction.

2. The method for predicting photovoltaic power of the day ahead according to claim 1, characterized in that: The historical photovoltaic power generation data is clustered based on similar days and divided into six sub-datasets: sunny days in winter, cloudy days in winter, rainy days in winter, sunny days in summer, cloudy days in summer, and rainy days in summer. Each training data set includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, and radiation intensity. The label data includes: historical power generation data.

3. The method for predicting photovoltaic power of the day ahead according to claim 2, characterized in that: In step S3, the hyperparameters of the photovoltaic power prediction model are optimized by the AOA algorithm, and the hyperparameters include: T, depth, η, λ, and num_leaves.

4. A day-ahead photovoltaic power prediction device, characterized in that: include: Acquisition module, used to obtain historical photovoltaic power generation data; The training module is used to train the CatBoost model based on historical photovoltaic power generation data to obtain a photovoltaic power prediction model; The optimization module is used to optimize the hyperparameters of the photovoltaic power prediction model to obtain an optimized photovoltaic power prediction model; wherein, real-time photovoltaic power generation data is input into the optimized photovoltaic power prediction model to perform day-ahead photovoltaic power prediction.

5. The day-ahead photovoltaic power prediction device according to claim 4, characterized in that: The acquisition module is also used to perform clustering processing on the historical photovoltaic power generation data based on similar days and divide it into six sub-datasets: sunny days in winter, cloudy days in winter, rainy days in winter, sunny days in summer, cloudy days in summer, and rainy days in summer. Each training data set includes: feature data and label data. The feature data includes: solar zenith angle, sunshine clarity, relative humidity, precipitation, pressure, wind speed, cloud type, and radiation intensity. The label data includes: historical power generation data.

6. The day-ahead photovoltaic power prediction device according to claim 5, characterized in that: The training module is used to optimize the hyperparameters of the photovoltaic power prediction model through the AOA algorithm, and the hyperparameters include: T, depth, η, λ, and num_leaves.

7. A day-ahead photovoltaic power prediction system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the day-ahead photovoltaic power prediction method according to any one of claims 1 to 3 is executed.

8. A storage medium, characterized in that: The storage medium stores a computer program, and the computer program executes the day-ahead photovoltaic power prediction method according to any one of claims 1 to 3 when running.