Model construction method for predicting photovoltaic generating capacity

By combining the spatial positional relationship and meteorological data of the target area and its neighboring areas, linear and random forest prediction models are trained, and the problem of low accuracy of photovoltaic power generation prediction in the existing technology is solved, achieving high accuracy and stability prediction effects.

CN120508923APending Publication Date: 2025-08-19XJ ELECTRIC CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510434432.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the prior art, the photovoltaic power generation prediction model constructed only considers weather factors and equipment factors, there is a problem of low prediction accuracy.

Method used

By training a linear prediction model and a random forest prediction model, combining the spatial position relationship and meteorological data of the target area and its neighboring areas, the parameters of the random forest prediction model are adjusted to generate a high-precision photovoltaic power generation prediction model.

Benefits of technology

High accuracy and high stability of photovoltaic power generation prediction can better understand the spatial distribution of power generation of photovoltaic equipment, provide a basis for photovoltaic enterprises and policy makers, and optimize the location selection of photovoltaic power stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508923A_ABST
    Figure CN120508923A_ABST
Patent Text Reader

Abstract

The invention relates to a model construction method for predicting photovoltaic power generation capacity, which comprises the following steps of: training a linear prediction model for predicting photovoltaic power generation according to a spatial position relationship between a to-be-measured target region and an adjacent region with photovoltaic power generation data, the photovoltaic power generation data of each region and meteorological data of the target region; according to the meteorological data and the photovoltaic power generation data of the target region, training a random forest prediction model used for predicting photovoltaic power generation of the target region; in the training process, according to the deviation between the power generation amount of the target area obtained by the linear prediction model and the power generation amount of the target area obtained by the random forest prediction model, the training parameters of the random forest prediction model are adjusted to complete the training of the random forest prediction model, and a prediction model with high precision and comprehensive analysis is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a model construction method for predicting photovoltaic power generation, and belongs to the technical field of power systems. Background Art

[0002] With the global transformation of energy structures and growing awareness of environmental protection, photovoltaic power generation, as a clean, renewable energy source, has garnered widespread attention and application. However, photovoltaic power generation is highly intermittent and volatile, influenced by a variety of factors, including weather, time of day, and environmental factors. This poses challenges to the stable operation of power grids and power dispatch. Therefore, accurately predicting photovoltaic power generation is crucial for improving the economic efficiency, stability, and reliability of power grids.

[0003] The published text of the Chinese invention patent application with application publication number CN117196103A discloses a distributed photovoltaic power prediction method and system based on numerical weather forecast, wherein distributed photovoltaic related data (including distributed photovoltaic geographical location information, numerical weather forecast historical data, distributed photovoltaic equipment data, distributed photovoltaic historical power data and distributed photovoltaic historical prediction data) are obtained; based on the numerical weather forecast historical data and the distributed photovoltaic historical power data, a prediction correlation feature set is generated; based on the numerical weather forecast historical data, the distributed photovoltaic equipment data and the distributed photovoltaic historical prediction error data, an error correlation feature set is generated; according to the error correlation feature set and several prediction methods, several error prediction models are constructed; and according to the prediction correlation feature set and the optimal error prediction model, a power value prediction model is constructed. Among them, historical weather forecast data includes irradiance, temperature, humidity and wind speed; distributed photovoltaic geographical location information includes longitude and latitude, and historical weather forecast data is obtained based on the geographical location information. At the same time, in the process of constructing the power value prediction model, historical weather forecast data, photovoltaic historical power data, photovoltaic equipment data and photovoltaic historical prediction errors are analyzed and processed. However, due to the diversity and complexity of the operating environment of photovoltaic equipment, there are many influencing factors in real life in addition to traditional weather-related factors. The prediction model constructed by simply considering weather factors and equipment factors has the problems of low prediction accuracy and poor effect. Summary of the Invention

[0004] The purpose of the present invention is to provide a model construction method for predicting photovoltaic power generation, so as to solve the problem of inaccurate prediction caused by the prediction model constructed by considering only weather factors and equipment factors.

[0005] To achieve the above objectives, the present invention proposes a model construction method for predicting photovoltaic power generation, comprising the following steps:

[0006] 1) training a linear prediction model for predicting photovoltaic power generation based on the spatial relationship between the target region to be measured and its neighboring regions with photovoltaic power generation data, the photovoltaic power generation data of each region, and the meteorological data of the target region; in the linear prediction model, the predicted power generation of the target region is obtained by weighted summation of the meteorological data of the target region and the photovoltaic power generation data of its neighboring regions, and the weight of the photovoltaic power generation data of the neighboring regions is determined based on the spatial relationship between the corresponding region and the target region;

[0007] 2) Based on the meteorological data and photovoltaic power generation data of the target area, a random forest prediction model is trained to predict the photovoltaic power generation of the target area. During the training process, the parameters of the random forest prediction model training are adjusted based on the deviation between the power generation of the target area obtained by the linear prediction model and the power generation of the target area obtained by the random forest prediction model to complete the training of the random forest prediction model.

[0008] Furthermore, the weights of the photovoltaic power generation data of the neighboring areas are obtained through the following steps:

[0009] Analyze the geographic data of the target area and its neighboring areas to obtain the spatial location relationship between the areas;

[0010] The photovoltaic power generation data of the target area and its neighboring areas were analyzed to obtain the spatial autocorrelation coefficient of the photovoltaic power generation of the neighboring areas to the photovoltaic power generation of the target area;

[0011] The weights of photovoltaic power generation data in neighboring areas are obtained through spatial position relationship and spatial autocorrelation coefficient.

[0012] Furthermore, the linear prediction model also includes geographical data of the target area. When the linear prediction model performs weighted summation, the geographical data of the target area, the meteorological data of the target area and the photovoltaic power generation data of its neighboring areas are also weighted summed.

[0013] Furthermore, the geographic data includes any one or two or more of longitude, latitude, slope and elevation.

[0014] Furthermore, the random forest prediction model for predicting photovoltaic power generation in the target area is trained through the following steps:

[0015] 1) Select target meteorological data with strong correlation with photovoltaic power generation data from all meteorological data containing several types of meteorological data;

[0016] 2) Constructing a random forest prediction model, and inputting the target class meteorological data and photovoltaic power generation data into the random forest prediction model for model training.

[0017] Furthermore, target meteorological data is selected by the following method:

[0018] The correlation coefficient between each type of meteorological data and photovoltaic power generation data is calculated respectively; and the meteorological data category whose correlation coefficient reaches a preset correlation coefficient threshold is determined as the target type of meteorological data.

[0019] Furthermore, the target meteorological data includes any one or two or more of solar radiation, temperature and humidity; wherein the solar radiation includes total radiation, direct radiation and scattered radiation.

[0020] The present invention has the following beneficial effects: a linear prediction model for predicting photovoltaic power generation is trained based on the spatial position relationship between a target region to be measured and its neighboring regions with photovoltaic power generation data, the photovoltaic power generation data of each region, and the meteorological data of the target region; in the linear prediction model, the predicted power generation of the target region is obtained by weighted summation of the meteorological data of the target region and the photovoltaic power generation data of its neighboring regions, and the weight of the photovoltaic power generation data of the neighboring regions is determined according to the spatial position relationship between the corresponding region and the target region; a random forest prediction model for predicting the photovoltaic power generation of the target region is trained based on the meteorological data and photovoltaic power generation data of the target region; during the training process, parameters of the random forest prediction model training are adjusted based on the deviation between the power generation of the target region obtained by the linear prediction model and the power generation of the target region obtained by the random forest prediction model to complete the training of the random forest prediction model, so as to realize the combination of the spatial position relationship between the target region and its neighboring regions in the process of predicting photovoltaic power generation, thereby predicting and analyzing photovoltaic power generation from the meteorological and spatial perspectives, and using the linear prediction model whose training data includes the spatial position relationship to intervene in the training process of the random forest prediction model, thereby generating a prediction model with high accuracy, high stability, and comprehensive analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a flow chart of a model construction method for predicting photovoltaic power generation proposed by the present invention;

[0022] Figure 2 This is a flow chart of a random forest prediction model in a practical application scenario of a model construction method for predicting photovoltaic power generation proposed in the present invention. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0024] The inventive concept of the present invention is: in order to deeply understand the impact of the operating environment factors and meteorological factors of photovoltaic equipment on its power generation, breaking through the traditional training mode of only considering meteorological factors in the training process of the random forest prediction model, integrating the spatial position relationship between the target area and its neighboring areas, so that the final random forest prediction model performs data analysis in the meteorological data dimension and the spatial data dimension, making the prediction results more accurate.

[0025] Method Example 1:

[0026] like Figure 1 FIG. 1 is a flow chart of a method for constructing a model for predicting photovoltaic power generation proposed by the present invention, which includes steps S11 and S12. Specifically:

[0027] Collect geographic data of the target area to be measured and its neighboring areas with photovoltaic power generation data, and analyze the spatial position relationship between each area based on the geographic data, wherein the geographic data includes any one or two or more of longitude, latitude, slope and elevation; the spatial position relationship refers to the spatial relationship between each area obtained by analyzing the geographic data of different areas, and the spatial relationship includes but is not limited to whether it is directly adjacent, the distance relationship, and the elevation difference relationship. The spatial relationship between different areas can be analyzed by using a spatial weight matrix or a spatial weight vector to represent the position relationship of different areas. In a preferred embodiment of the present application, the spatial position relationship between the target area and its neighboring areas is preferably represented by a spatial weight matrix, wherein the weight value corresponding to the spatial relationship can be expressed as: when the spatial When the relationship is whether it is a direct border relationship, the weight value of direct border is 1, and the weight value of indirect border is 0; when the spatial relationship is a distance relationship, the weight value is inversely proportional to the distance difference, and the weight value can be obtained by the inverse distance square function; when the spatial relationship is an elevation difference, the weight value is set to 0 when the elevation difference exceeds the threshold; the neighboring area with photovoltaic power generation data refers to an area that directly or indirectly borders the target area and has photovoltaic power generation data, which can be an area within a set range. For example, the area within 10km of the target area and with photovoltaic power generation data is set as a neighboring area; for another example, the area with photovoltaic power generation data that can be reached within 5 hours' drive around the target area is set as a neighboring area. The specific selection and adjustment of the neighboring area will be based on the distribution of photovoltaic equipment or the prediction range.

[0028] After obtaining the spatial position relationship between each region, step S11 is executed to train a linear prediction model for predicting photovoltaic power generation based on the spatial position relationship between the target region to be measured and its neighboring regions with photovoltaic power generation data, the photovoltaic power generation data of each region, and the meteorological data of the target region. In the linear prediction model, the predicted power generation of the target region is obtained by weighted summation of the meteorological data of the target region and the photovoltaic power generation data of its neighboring regions, and the weight of the photovoltaic power generation data of its neighboring regions is determined according to the spatial position relationship between the corresponding region and the target region. Here, the photovoltaic power generation data includes but is not limited to power generation, power generation, etc.; the meteorological data includes but is not limited to common weather-related data such as solar radiation data, temperature data, humidity data, wind speed data, wind direction data, and air pressure data.

[0029] Step S12, based on the meteorological data and photovoltaic power generation data of the target area, train a random forest prediction model for predicting the photovoltaic power generation of the target area; during the training process, according to the deviation between the power generation of the target area obtained by the linear prediction model and the power generation of the target area obtained by the random forest prediction model, adjust the parameters of the random forest prediction model training to complete the training of the random forest prediction model; here, when the deviation value reaches the predicted deviation threshold used to judge the high level of consistency between the prediction results of the models, the training of the random forest prediction model is completed; in actual application, the adjusted parameters include but are not limited to the number of decision trees, the depth of the decision tree, the minimum number of samples for splitting nodes, the minimum number of samples for leaf nodes, the number of features considered for each split, and the sample subsampling ratio.

[0030] Different adjustment strategies correspond to different adjustment parameters. Specifically, when adjusting the number of decision trees, if the deviation is larger, it means that the random forest prediction model is significantly different from the linear prediction model. In this case, the number of decision trees should be increased (for example, from 100 to 300) to enhance the model capacity and adjust the prediction results. If the deviation is close to the threshold but the training time is too long, the number of decision trees should be reduced to balance accuracy and efficiency.

[0031] When adjusting the depth of the decision tree, if the deviation is large and the model is underfitting, it means that the random forest prediction model is too conservative. In this case, the maximum depth can be increased (e.g., from 5 to 15) to allow the decision tree to generate more fine-grained rules. If the deviation is small but the validation set performance is poor, it means that the random forest prediction model may be overfitting. In this case, the maximum depth can be limited (e.g., from 20 to 10) to reduce the complexity of the model.

[0032] When adjusting the minimum number of samples for splitting a node, if the deviation is large and the random forest prediction model is too coarse, reduce the value (for example, from 20 to 5) to allow the node to split more times; if the deviation is small but the random forest prediction model is sensitive to the training data, increase the value (for example, from 5 to 50) to prevent overfitting to noise.

[0033] When adjusting the minimum number of samples for leaf nodes, if the deviation is large and the random forest prediction model ignores local features, reduce the value (for example, from 10 to 1) to allow finer-grained leaf nodes; if the deviation is small but the random forest prediction model has large fluctuations in predictions, increase the value (for example, from 1 to 20) to smooth the prediction results.

[0034] When adjusting the number of features considered for each split, if the bias is large and the interactions between features are not fully captured, increase the number of features (for example, change from sqrt(n_features) to log2(n_features)); if the bias is small but the variance of the random forest prediction model is high, reduce the number of features (for example, change from all to sqrt(n_features)) to improve generalization ability.

[0035] When adjusting the sample subsampling ratio, if the deviation is large and the random forest prediction model does not fully utilize the data, increase the subsampling ratio (for example, from 0.8 to 1.0); if the deviation is small but the random forest prediction model overfits the training set, reduce the subsampling ratio (for example, from 1.0 to 0.6) to introduce more randomness.

[0036] Through steps S11 and S12, the spatial position relationship between different regions is introduced into the training process of the random forest prediction model, multi-dimensional model training is achieved, and the prediction accuracy of the random forest prediction model is improved. This avoids the problem of large deviation and inaccuracy in the prediction results of photovoltaic power generation when only weather-related factors are considered in the diverse and complex operating environment of photovoltaic equipment.

[0037] In a preferred embodiment of the invention, a target area A is preferred, which corresponds to neighboring areas B and C with photovoltaic power generation data; geographic data (preferably longitude data, latitude data and elevation data) of the target area A and its corresponding neighboring areas B and C are collected; the geographic data of all areas are analyzed to generate a spatial weight matrix W for the three areas; based on the spatial weight matrix W, the photovoltaic power generation of the three areas (that is, the photovoltaic power generation data is preferably photovoltaic power generation) and the solar radiation data, temperature data and humidity data of the target area A (that is, the meteorological data preferably includes only solar radiation data, temperature data and humidity data). A linear prediction model model1 is generated.

[0038] Based on the solar radiation data, temperature data, humidity data and photovoltaic power generation of the target area A, the random forest prediction model model2 is trained. During the training process of model2, it is preferred to output the prediction result y1 by calling model1 and output the prediction result y2 by calling the model2 that is still being trained; the deviation value between y1 and y2 is calculated, and the training of model2 is adjusted according to the deviation value until the prediction results of model1 are highly consistent with the prediction results of model2, and the training of model2 is completed.

[0039] Method Example 2:

[0040] The present invention proposes a model construction method for predicting photovoltaic power generation. In step S11, the weights of photovoltaic power generation data of neighboring areas are obtained through the following steps:

[0041] The geographic data of the target area and its neighboring areas are analyzed to obtain the spatial position relationship between the areas, wherein the geographic data include any one or two or more of longitude, latitude, slope and elevation. For example, when the geographic data includes longitude, the longitude difference relationship between the areas is obtained based on the comparison of the longitudes between the areas, and the longitude difference relationship between different areas is analyzed. The longitude difference relationship between different areas can be represented by a spatial weight matrix; when the geographic data includes longitude and latitude, the distance relationship between the areas is obtained based on the comparison of the longitude and longitude between the areas, and the distance relationship between different areas can be represented by a spatial weight matrix. When the geographic data includes longitude, latitude and slope, the distance and altitude relationship between the areas is obtained based on the comparison of the longitude, latitude and slope between the areas, and the distance and altitude relationship between different areas is analyzed. The distance and altitude relationship between different areas can be represented by a spatial weight matrix.

[0042] The photovoltaic power generation data of the target area and its neighboring areas were analyzed to obtain the spatial autocorrelation coefficient of the photovoltaic power generation of the neighboring areas to the photovoltaic power generation of the target area. For example, the photovoltaic power generation of the target area and the neighboring areas within 100 km outside the target area was analyzed, and the spatial autocorrelation coefficient of the photovoltaic power generation of the neighboring areas within 1-50 km outside the target area to the photovoltaic power generation of the target area was 0.2, and the spatial autocorrelation coefficient of the photovoltaic power generation of the neighboring areas within 50-100 km outside the target area to the photovoltaic power generation of the target area was 0.1.

[0043] The weight of the photovoltaic power generation data of the neighboring areas is obtained through the spatial position relationship and the spatial autocorrelation coefficient. Here, when the spatial position relationship is represented by the spatial weight matrix, the weight of the photovoltaic power generation data of the target neighboring areas is represented by calculating the product between the spatial weight matrix and the spatial autocorrelation coefficient of the target neighboring areas.

[0044] Following the above embodiment of the present invention, the geographical data of the target area A and its corresponding adjacent areas B and C are collected and analyzed to generate the spatial weight matrix W of the three areas; the photovoltaic power generation of the three areas is analyzed, and the spatial autocorrelation coefficient of the adjacent area B to the target area A is obtained as ρ B ; The spatial autocorrelation coefficient of the neighboring area C to the target area A is ρ C ; ρ B *W is set as the weight of the photovoltaic power generation in the neighboring area B; ρ C *W is set as the weight of the photovoltaic power generation of the neighboring region C, which is used for training the linear prediction model.

[0045] Method Example 3:

[0046] The present invention proposes a model construction method for predicting photovoltaic power generation. In step S11, the linear prediction model also includes geographic data of the target area. When the linear prediction model performs weighted summation, the target area's geographic data is also weighted summed together with the target area's meteorological data and the photovoltaic power generation data of its neighboring areas. The geographic data includes any one, two, or more of longitude, latitude, slope, and elevation. Preferably, the geographic data includes longitude, latitude, slope, and elevation, and the linear prediction model can be expressed by the following formula:

[0047]

[0048] Where Y is the predicted photovoltaic power generation; are 1-m independent variables that affect photovoltaic power generation, where β i is the coefficient of the ith independent variable, X i is the measured value of the ith independent variable; ρ is the spatial autocorrelation coefficient; W is the spatial weight matrix; Y2 is the photovoltaic power generation vector of the neighboring area; ε is the random error; β0 is the constant term;

[0049] Among them, the independent variable refers to a certain type of meteorological data in the meteorological data of the target area and a certain type of geographical data in the geographical data; it should be noted that the meteorological data includes but is not limited to common weather-related data such as solar radiation data, temperature data, humidity data, wind speed data, wind direction data and air pressure data; any one or two or more of the longitude, latitude, slope and elevation of the geographical data, that is, the geographical data can be divided into longitude data, latitude data, slope data and elevation data; a certain type of meteorological data in the meteorological data is used as an independent variable to represent the impact of this type of meteorological data on photovoltaic power generation; similarly, a certain type of geographical data in the geographical data is used as an independent variable to represent the impact of this type of geographical data on photovoltaic power generation.

[0050] For example, considering the impact of multiple factors on photovoltaic power generation, data on photovoltaic power generation, geographic factors (latitude and longitude, topography), and environmental factors (solar radiation, temperature, and humidity) are collected for the target region and its neighboring areas. Based on this data, the dependent and independent variables of the model are defined. The dependent variable (Y) is photovoltaic power generation (kilowatt-hours generated per hour or day), and the independent variables (X) are X1: solar radiation intensity (e.g., watts per square meter); X2: ambient temperature (degrees Celsius or Fahrenheit); X3: relative humidity (percentage); X4: longitude; X5: latitude; X6: slope; and X7: elevation. To account for spatial autocorrelation and spatial heterogeneity, a spatial lag model (SLM) is used in the spatial regression model. Specifically, it is expressed as follows: Y = β0 + β1*X1 + β2*X2 + β3*X3 + β4*X4 + β5*X5 + β6*X6 + β7*X7 + ρ*WY2 + ε.

[0051] Where, Y: predicted photovoltaic power generation, i.e., the dependent variable.

[0052] β0: intercept term, also known as constant term, which represents the photovoltaic power generation when all independent variables are zero.

[0053] β1*X1: The impact of solar radiation on photovoltaic power generation. β1 is the coefficient of solar radiation, which indicates the change in photovoltaic power generation per unit change in solar radiation. X1 is the measured value of solar radiation.

[0054] β2*X2: The effect of temperature on photovoltaic power generation. β2 is the temperature coefficient, which indicates the change in photovoltaic power generation per unit change in temperature. X2 is the measured value of temperature.

[0055] β3*X3: The impact of humidity on photovoltaic power generation. β3 is the coefficient of humidity, which indicates the change in photovoltaic power generation per unit change in humidity. X3 is the measured value of humidity.

[0056] β4*X4: The impact of longitude on photovoltaic power generation. β4 is the coefficient of longitude, which represents the change in photovoltaic power generation per unit change in longitude. X4 is the measured value of longitude.

[0057] β5*X5: The impact of latitude on photovoltaic power generation. β5 is the coefficient of latitude, which represents the change in photovoltaic power generation per unit change in latitude. X5 is the measured value of latitude.

[0058] β6*X6: The impact of slope on photovoltaic power generation. β6 is the slope coefficient, which represents the change in photovoltaic power generation per unit change in slope. X6 is the measured value of the slope.

[0059] β7*X7: The impact of elevation on photovoltaic power generation. β7 is the coefficient of elevation, which indicates the change in photovoltaic power generation per unit change in elevation. X7 is the measured value of elevation.

[0060] ρWY2: Spatial autocorrelation term. ρ is the spatial autocorrelation coefficient, which represents the influence of PV power generation in neighboring regions on PV power generation in the target region. W is the spatial weight matrix, which represents the spatial relationship between different regions. Y2 is the PV power generation vector of neighboring regions. This term indicates that PV power generation in the target region is affected not only by the conditions in the target region but also by the PV power generation in neighboring regions. The spatial weight matrix W is derived from the geographic data of the target region and its neighboring regions; the spatial autocorrelation coefficient ρ is derived from the PV power generation data of the target region and its neighboring regions.

[0061] ∈: Random error term, which is the variation not explained in the model, that is, the difference between the actual observations and the values predicted by the model.

[0062] Using statistical software (such as R, Python, etc.), the photovoltaic power generation of neighboring areas, the photovoltaic power generation of the target area, solar radiation intensity data; ambient temperature data (Celsius or Fahrenheit); relative humidity data (percentage); longitude data; latitude data; slope data; and elevation data are all input into the above formula to train the parameters in the model (i.e., β0-β7, ρ, ∈) to obtain the weights of each type of data; the predictive ability of Y is verified through cross-validation or other statistical tests; Y is used to predict the photovoltaic power generation of new or currently unobservable target areas, so that geographical factors, environmental factors, and spatial autocorrelation factors of photovoltaic power generation in neighboring areas are taken into account in the process of constructing the prediction model. This can not only help photovoltaic power generation companies and policymakers better understand the spatial distribution of the current power generation of photovoltaic equipment, but also provide a basis for the site selection and optimization of photovoltaic power stations.

[0063] Method Example 4:

[0064] The present invention proposes a model construction method for predicting photovoltaic power generation. In step S12, a random forest prediction model for predicting photovoltaic power generation in a target area is trained by the following steps:

[0065] 1) Select target meteorological data with a strong correlation with photovoltaic power generation data from all meteorological data including several types of meteorological data; here, the strength of the correlation between the several types of meteorological data and the photovoltaic power generation data can be represented by a correlation coefficient; it can also be represented by a heat map; it can also be represented by a combination of the correlation coefficient and the heat map, which is not limited in the present invention. At the same time, in actual application scenarios, the correlation between various types of meteorological data is also considered, and then the target meteorological data with a strong correlation is comprehensively judged and selected.

[0066] When using the correlation coefficient to represent the strength of the correlation between several types of meteorological data and photovoltaic power generation data, calculate the correlation coefficient between each type of meteorological data and photovoltaic power generation data respectively; determine the category of meteorological data whose correlation coefficient reaches the preset correlation coefficient threshold as the target category of meteorological data; the correlation coefficient can be calculated through the pearson correlation coefficient formula. Specifically, the pearson correlation coefficient calculation formula is as follows:

[0067]

[0068] where, x i is a certain type of meteorological data corresponding to the i-th moment; is the mean value of this type of meteorological data; y i is the photovoltaic power generation power at the i-th moment; is the average value of photovoltaic power generation power.

[0069] For example, through the pearson correlation coefficient calculation formula, calculate the correlation coefficients between solar radiation data, temperature data, humidity data, wind speed data, wind direction data, and air pressure data and photovoltaic discharge power respectively. According to the analysis of all correlation coefficients, delete the three variables with weak correlations, namely wind speed data, wind direction data, and air pressure data, and select solar radiation data, temperature data, and humidity data as the target category of correlation coefficients. And, among them, the solar radiation category includes total irradiance category, direct irradiance category, and diffuse irradiance category. That is, use the five influencing factors of total irradiance, direct irradiance, diffuse irradiance, temperature category, and humidity category as the input of the random forest prediction model.

[0070] 2) Build a random forest prediction model, and input the target category of meteorological data and photovoltaic power generation data into the random forest prediction model for model training. The training process of the random forest prediction model, specifically speaking, assume that the size of the training data T is A, the number of features is B, and the size of the random forest is C. The specific steps of the random forest algorithm are as follows: First, traverse the size C of the random forest: Sampling A times from the training data T with replacement to form a new sub-training data E; randomly select b features, where b < B; use the new training data E and b features to learn a complete decision tree. The model is trained through C rounds to form a sequence of base estimators {t1(X), t2(X), ……, t C (X)}, and then average their prediction results or use the majority voting principle to obtain the final prediction result of the integrated estimator. The final prediction result can be expressed by the following formula:

[0071]

[0072] Among them, H(x) is the final classification result of the model; F is the indicator function; t i is a single decision tree classifier (base evaluator); Y is the output variable (target variable).

[0073] During the training process, the following formula can be used to optimize the training of the generalization error upper bound:

[0074]

[0075] Where G is the model generalization error; is the correlation between decision trees; Q is the classification strength of the decision tree.

[0076] Method Example 5:

[0077] like Figure 2 The figure shows a flow chart of a random forest prediction model in an actual application scenario in a model construction method for predicting photovoltaic power generation proposed by the present invention, wherein, first, the collected data is preprocessed, that is, converted into a processable data type, specifically, redundant data is cleaned and missing data is processed. The preprocessed data is further compared with the data required for modeling, mainly to check the proportion of missing values in the preprocessed data, the data type of each data column, whether there is inconsistency in the format of some data, etc. For the problem of missing values, the interfering data is deleted, and the missing values are filled using linear interpolation and the mean of the previous and next values, and then each column is processed using a fitting curve interpolation; for the problem of abnormal data, two methods of setting breakpoints and replacing the average value are used respectively; finally, the standardization is performed using the following formula:

[0078]

[0079] Among them, x` is the formatted standard data; x is the original data that has not been standardized; x max 、x min are the maximum and minimum values of the original data before normalization.

[0080] Next, the correlation coefficients between various meteorological data and the correlation coefficients between various meteorological data and actual photovoltaic output power (i.e., photovoltaic power generation data) were calculated from the standardized data for meteorological data, and a heat map of the correlation coefficients was obtained. A comprehensive analysis was performed based on the heat map, and finally the three variables with weak correlation, wind speed, wind direction, and air pressure, were deleted. Five influencing factors, namely total irradiance, direct irradiance, diffuse irradiance, temperature, and humidity, were selected as training data input.

[0081] Next, a random forest prediction model is established and training data is input for training. During the training process, the prediction results of the linear prediction model and the upper bound formula of the generalization error are continuously used to optimize the model and determine whether the evaluation results are optimal. If not, the parameters of the random forest prediction model in training are optimized and the model continues to be trained. If so, the optimal random forest prediction model for predicting photovoltaic power generation in the target area is generated.

[0082] Finally, the weather forecast for the next three days is input into the random forest prediction model to obtain the power forecast results for the next three days.

[0083] In summary, the present invention highlights the impact of weather and geographical factors on photovoltaic power generation, combines the advantages of multiple prediction models, and achieves a more comprehensive and accurate prediction of photovoltaic power generation, thereby improving the accuracy of the prediction, which is of great significance for promoting the development of the new energy industry and promoting the stable operation of the power system.

[0084] The prediction model finally obtained through training has the characteristics of high precision (through feature analysis and ensemble learning, it can make full use of data information of multiple influencing factors to improve the accuracy of prediction), high stability (ensemble learning can reduce the instability of the prediction results of a single model by combining multiple prediction models) and strong adaptability (with the accumulation of data and the updating of the model, the prediction model can continuously adapt to new environments and conditions, improving the adaptability of the prediction).

Claims

1. A model construction method for predicting photovoltaic power generation, characterized in that: The steps include: 1) Training a linear prediction model for predicting photovoltaic power generation based on the spatial relationship between the target region and its neighboring regions with photovoltaic power generation data, the photovoltaic power generation data of each region, and the meteorological data of the target region; in the linear prediction model, the predicted power generation of the target region is obtained by the weighted summation of the target region's meteorological data and the photovoltaic power generation data of its neighboring regions, with the weight of the photovoltaic power generation data of the neighboring regions being determined based on the spatial relationship between the corresponding region and the target region; 2) Based on the meteorological data and photovoltaic power generation data of the target area, a random forest prediction model is trained to predict the photovoltaic power generation in the target area. During the training process, the parameters of the random forest prediction model training are adjusted based on the deviation between the power generation in the target area obtained by the linear prediction model and the power generation in the target area obtained by the random forest prediction model to complete the training of the random forest prediction model.

2. The model construction method for predicting photovoltaic power generation according to claim 1, characterized in that: The weights of the photovoltaic power generation data of the neighboring areas are obtained through the following steps: Analyze the geographic data of the target area and its neighboring areas to obtain the spatial location relationship between the areas; The photovoltaic power generation data of the target area and its neighboring areas were analyzed to obtain the spatial autocorrelation coefficient of the photovoltaic power generation of the neighboring areas to the photovoltaic power generation of the target area; The weights of photovoltaic power generation data in neighboring areas are obtained through spatial position relationship and spatial autocorrelation coefficient.

3. The model construction method for predicting photovoltaic power generation according to claim 1, characterized in that: The linear prediction model also includes geographical data of the target area. When the linear prediction model performs weighted summation, the geographical data of the target area, the meteorological data of the target area and the photovoltaic power generation data of its neighboring areas are also weighted summed.

4. The model construction method for predicting photovoltaic power generation according to claim 2 or 3, characterized in that: The geographic data includes any one or two or more of longitude, latitude, slope and elevation.

5. The model construction method for predicting photovoltaic power generation according to claim 1, characterized in that: The following steps are used to train a random forest prediction model for predicting photovoltaic power generation in the target area: 1) Select target meteorological data with strong correlation with photovoltaic power generation data from all meteorological data containing several types of meteorological data; 2) Constructing a random forest prediction model, and inputting the target class meteorological data and photovoltaic power generation data into the random forest prediction model for model training.

6. The model construction method for predicting photovoltaic power generation according to claim 5, characterized in that: Select the target meteorological data by the following method: The correlation coefficient between each type of meteorological data and photovoltaic power generation data is calculated respectively; and the meteorological data category whose correlation coefficient reaches a preset correlation coefficient threshold is determined as the target type of meteorological data.

7. The model construction method for predicting photovoltaic power generation according to claim 6, characterized in that: The target meteorological data includes any one or two or more of solar radiation, temperature and humidity; wherein the solar radiation includes total radiation, direct radiation and scattered radiation.

Citation Information

Patent Citations

  • Distributed photovoltaic power prediction method and system based on numerical weather forecast

    CN117196103A