A Whole-County Photovoltaic Forecasting Method Based on Cluster Partitioning and Data Augmentation

By clustering and data augmenting of distributed photovoltaics in the county, a photovoltaic prediction neural network model is built, which solves the problems of lack of distributed photovoltaic data and massive site prediction, and improves the accuracy of photovoltaic power generation prediction.

CN114492941BActive Publication Date: 2025-05-27SOUTHEAST UNIV +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111646119.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-27
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The distributed photovoltaics in the whole county have problems with lack of data and the prediction of massive sites, resulting in insufficient accuracy of photovoltaic power generation prediction.

Method used

Using a method based on cluster division and data augmentation, photovoltaic clusters are divided through DBSCAN clustering, photovoltaic data augmentation neural network model is constructed for data augmentation, and combined with CNN neural network for training to generate photovoltaic prediction neural network model.

Benefits of technology

The accuracy of photovoltaic power generation power prediction in the whole county has been improved, the distribution of numerical weather forecasts has been reduced, the calculation pressure has been reduced, and the prediction capability of the model has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492941B_ABST
    Figure CN114492941B_ABST
Patent Text Reader

Abstract

A method for predicting the distributed photovoltaic power of an entire county based on cluster division and data augmentation, specifically: select the typical power curve of sunny days from the historical database of photovoltaic power output of the entire county, and normalize the power output with the maximum power of a single station; calculate the Pearson correlation coefficient as the distance metric, and use the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm to cluster photovoltaic sites to form cluster divisions. For outliers, use k-nearest neighbor search and divide them into the nearest cluster; within the cluster, use a generative adversarial neural network to augment historical data; jointly train a deep convolutional network prediction model with the original data and the generated data in a pictorial form. The prediction method of the present invention learns the original data distribution through the dynamic game process training of the improved GAN, and then generates data with the corresponding distribution, supplementing the historical database of the distributed photovoltaic power of the entire county. By training the deep convolutional neural network with the enhanced training set, the prediction accuracy of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of photovoltaic power prediction, and particularly relates to a whole-county photovoltaic prediction method based on cluster division and data augmentation. Background Art

[0002] Climate change is a worldwide problem faced by mankind. With the rapid development of productivity, the amount of carbon dioxide emissions by various countries has increased sharply, and the greenhouse effect has become increasingly serious. In order to address climate change, countries around the world have set carbon emission reduction targets in the form of a global agreement, and China has also put forward the goals of carbon peak and carbon neutrality. Photovoltaic power generation does not require the combustion of fossil fuels for energy conversion and is a green and clean energy source. Concentrated photovoltaics require a large amount of land resources and are mostly concentrated in the sparsely populated northwest region. The electric power resources need to be transmitted to the load center through ultra-high voltage transmission lines. Rooftop distributed photovoltaics do not occupy dedicated land resources and are close to the load center, with huge power generation potential. As of the end of May 2021, 1,884,700 low-voltage distributed photovoltaic grid connections with a capacity of 45.3177 million kilowatts had been put into operation in the operating area of the State Grid Corporation. On September 14, 2021, the National Energy Administration officially issued the "Notice on Announcing the Pilot List of Rooftop Distributed Photovoltaic Development in Whole Counties (Cities, Districts) (Guonengzongtong Xinneng

[2021] No. 84)" to actively promote the grid connection work of whole-county photovoltaic users. The penetration rate of photovoltaics is continuously increasing, and the randomness and volatility of its output pose pressure on the safe and stable operation of the power grid. Therefore, the accuracy of whole-county photovoltaic power prediction is becoming increasingly important.

[0003] There are currently two main problems with whole-county distributed photovoltaics. One is that distributed photovoltaics started later than concentrated photovoltaics, and historical meteorological data and power generation power data are relatively scarce, making it difficult to train a high-precision data-driven model. The other is that there are a large number of distributed photovoltaics with a wide geographical distribution, making it difficult to perform one-site-one-prediction similar to concentrated sites. Therefore, it is necessary to solve the problem of model training with poor data, and at the same time, it is necessary to handle the whole-county photovoltaic power prediction problem of a large number of sites. Summary of the Invention

[0004] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a whole-county photovoltaic prediction method based on cluster division and data augmentation.

[0005] The present invention adopts the following technical solutions:

[0006] A whole-county photovoltaic prediction method based on cluster division and data augmentation, comprising the following steps:

[0007] S1, collect the historical sunny-day photovoltaic output data and meteorological data of each photovoltaic output site within the set area and within the set time;

[0008] S2. Calculate the maximum power of each PV output site from the sunny-day PV output data collected in S1, perform normalization, and then divide the PV clusters.

[0009] S3. Construct a PV data augmentation neural network model.

[0010] S4. Perform data augmentation within the PV clusters using the PV data augmentation neural network model constructed in S3, and then input the augmented data into the CNN neural network. After training, obtain the PV prediction neural network model.

[0011] S5: Input the meteorological forecast data into the PV prediction neural network model in S4 for PV prediction.

[0012] In S1, the PV output data is the historical sunny-day PV output data of the whole county.

[0013] Meteorological data includes month, day, hour, minute, direct irradiance, diffuse horizontal irradiance, total horizontal irradiance, ambient temperature, air pressure, relative humidity, wind direction, wind speed, surface reflectance, and power generation.

[0014] In S2, the historical sunny-day PV output curve is a normalized output curve. First, calculate the maximum power of a single station from the historical output data, and then calculate the normalized value of the curve.

[0015] The formula for data normalization is as follows:

[0016]

[0017] In the above formula, x is the original sample data, x max is the maximum historical output value of the site, and z is the normalized data.

[0018] In S2, use the DBSCAN method to divide the PV clusters. The distance metric of DBSCAN uses the pearson correlation distance. For the abnormal sites generated by DBSCAN for clustering analysis, use the k-nearest neighbor algorithm to search for the nearest k sites with the pearson correlation distance as the metric, and count the cluster numbers of the k sites. Divide the abnormal sites into the cluster with the largest number of numbers.

[0019] In S3, the optional neural network model is GAN.

[0020] The optional neural network model can also be an improved GAN. The specific algorithm is as follows:

[0021] S3.1. Set the learning rate α, truncation parameter c, number of batch training samples m, and number of times the discriminator iterates for each iteration of the generator n critic , initialize the discriminator network parameter w t , generator network parameter θt ;

[0022] S3.2, Check the generator parameter θ t to see if it converges. If it converges, end the iteration; if not, go to S3.3;

[0023] S3.3, Check if the current iteration number reaches the iteration number threshold n critic . If not, update the discriminator network parameter w t and update the current iteration number, and repeat this step; if the iteration number threshold n critic is reached, go to step S3.4;

[0024] Step 3.4, Calculate the generator loss function and update the generator parameter θ t , and return to step 3.2;

[0025] The method to update the discriminator network parameter w t is as follows:

[0026]

[0027] where clip(.) represents the clipping function in deep learning, and RMSProp(.) represents the learning rate adaptive optimizer in deep learning, represents the gradient of the discriminator loss function, and the specific calculation method is:

[0028]

[0029] where x (i) is the i-th training batch in the input parameters; z (i) is the batch sampling from the sample distribution generated by the generator for the i-th training; is the discriminator network.

[0030] The gradient of the generator loss function is:

[0031]

[0032] where is the discriminator network, and z (i) is the batch sampling from the sample distribution generated by the generator for the i-th training.

[0033] The method to update the generator parameter θ t is as follows:

[0034]

[0035] where BMSProp(.) represents the learning rate adaptive optimizer in deep learning.

[0036] The data augmentation of S4 means that first, the data collected by S1 is converted into pictures. The original data of one day is divided by day, with 48 points per day, forming pictures with a size of 48 * 15 pixels. Then, through the photovoltaic data augmentation neural network model, adversarial learning is carried out to generate new data with the same distribution.

[0037] In S4, the CNN neural network has a 5-layer structure. Each of the first four layers contains a convolutional layer, a batch normalization layer, and a rectified linear unit layer. The last layer is a fully connected layer, a batch normalization layer, and a ReLU layer.

[0038] The beneficial effects of the present invention are that, compared with the prior art,

[0039] 1. The prediction method of the present invention divides the distributed photovoltaics in the whole county into clusters, divides the distributed sites with strong correlation into one cluster, and predicts the overall power output of the photovoltaic cluster, so as to reduce the numerical weather prediction points, effectively reduce the computational pressure, and improve the overall prediction accuracy;

[0040] 2. The prediction method of the present invention learns the mutual relationship and distribution characteristics of the original data through adversarial training, generates new similar data, supplements the poor training data set, and improves the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the flow chart of the prediction method of the present invention;

[0042] Figure 2 is the bar chart of DBSCAN clustering of the present invention;

[0043] Figure 3 is the time series chart of the DBSCAN clustering result of the present invention;

[0044] Figure 4 is the result after k-nearest neighbor search of the present invention;

[0045] Figure 5 is the visualization of the original data;

[0046] Figure 6 is the data generated by the improved GAN and the training process;

[0047] Figure 7 is the comparison chart of the prediction time series of CNN and the improved GAN-CNN. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following further describes the present application with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present application.

[0049] A whole - county photovoltaic prediction method based on cluster division and data augmentation, and the flow of the prediction method is as follows Figure 1 shown, and specifically includes the following steps:

[0050] S1: Collect the historical sunny - day photovoltaic output data and meteorological data of each photovoltaic output site within the set area and within the set time;

[0051] In the present invention, the collected photovoltaic output data is the historical sunny - day photovoltaic output data of the whole county;

[0052] The environmental data includes month, day, hour, minute, direct irradiance, diffuse horizontal irradiance, total horizontal irradiance, environmental temperature, air pressure, relative humidity, wind direction, wind speed, surface reflectance, and power generation.

[0053] S2: Calculate the maximum power of each photovoltaic output site based on the sunny - day photovoltaic output data collected in S1, perform normalization, and then conduct cluster division;

[0054] Furthermore, the historical sunny - day photovoltaic output curve in S2 is a normalized output curve. First, calculate the maximum power of a single station from the historical output data, and then calculate the normalized value of the curve;

[0055] The formula for data normalization is as follows:

[0056]

[0057] In the above formula, x is the original sample data, x max is the maximum historical output value of the site, and z is the normalized data.

[0058] According to the typical historical sunny - day photovoltaic output curve of the site, use DBSCAN to conduct cluster division on the whole - county photovoltaic to form a highly correlated photovoltaic cluster;

[0059] Those of ordinary skill in the art can select the method of cluster division according to the actual situation. The method selected in the present invention is only a preferred embodiment and cannot be an inevitable limitation of the present invention.

[0060] Furthermore, in S2, DBSCAN is used for cluster division, and the distance metric of DBSCAN uses the Pearson correlation distance;

[0061] The formula for the Pearson correlation distance is as follows:

[0062]

[0063] Among them,

[0064]

[0065] xs , x t is the output sequence of any two random sites, x sj denotes s the j-th data value in x ti denotes t the i-th data value in x, n denotes sj the total number of data in x ti the total number of data in x;

[0066] Furthermore, for the abnormal sites generated by DBSCAN clustering analysis, the k-nearest neighbor algorithm is used to search for the nearest k sites with the Pearson correlation distance as the metric, the cluster numbers of the k sites are counted, and the abnormal sites are divided into the cluster with the largest number of numbers.

[0067] S3: Construct a photovoltaic data augmentation neural network model;

[0068] Those of ordinary skill in the art can select a neural network according to the actual situation. The method recommended by the present invention is only a preferred embodiment and cannot be an inevitable limitation of the present invention.

[0069] Preferably, the selectable neural network model is GAN;

[0070] In this embodiment, the photovoltaic data augmentation neural network model is an improved GAN, and the algorithm of the improved GAN is specifically:

[0071] S3.1, set the learning rate α, truncation parameter c, number of batch training samples m, and number of times n that the discriminator iterates each time the generator iterates critic, Initialize the discriminator network parameter w t and the generator network parameter θ t ,

[0072] S3.2, check whether the generator parameter θ t converges. If it converges, end the iteration; if not, enter S3.3;

[0073] S3.3, check whether the current iteration number reaches the iteration number threshold n critic , if not, update the discriminator network parameter w t and update the current iteration number, and repeat this step; if the iteration number threshold n is reached, enter step S3.4;

[0074] Those of ordinary skill in the art can set the update method of the discriminator network parameter w t according to the actual situation. The method given in the present invention is only a preferred embodiment and cannot be an inevitable limitation of the present invention.

[0075] Update the discriminator network parameter wt The method is as follows:

[0076]

[0077] Among them, clip(.) represents the clipping function in deep learning, and RMSProp(.) represents the learning rate adaptive optimizer in deep learning. represents the gradient function of the discriminator loss function, and the specific calculation method is as follows:

[0078]

[0079] Among them, x (i) is the i-th training batch in the input parameters; z (i) is the batch sampling from the sample distribution generated by the generator in the i-th training; is a discriminator network, and those skilled in the art can select it according to the actual situation;

[0080] Step 3.4, calculate the generator loss function and update the generator parameters θ t , and return to Step 3.2;

[0081] Those of ordinary skill in the art can set the generator loss function and the update method of the generator parameters θ t The method given in the present invention is only a preferred embodiment and cannot be an inevitable limitation of the present invention.

[0082] The gradient of the generator loss function is:

[0083]

[0084] Update the generator parameters θ t The method is as follows:

[0085]

[0086] S4: Perform data augmentation through the photovoltaic data augmentation neural network model constructed in S3 within the photovoltaic cluster, and then input the augmented data into the CNN neural network to obtain the photovoltaic prediction neural network model after training.

[0087] Furthermore, the data collected in S1 is made into pictures, and the data of one day is divided by day, with 48 points per day, forming pictures with a size of 48 * 15 pixels; then, through the photovoltaic data augmentation neural network model, adversarial learning is performed to generate new data with the same distribution.

[0088] Furthermore, the CNN neural network has a 5-layer structure. Each of the first four layers contains a convolutional layer, a batch normalization layer, and a Rectified Linear Unit layer. The last layer is a fully connected layer, a batch normalization layer, and a ReLU layer. The original data and new data are used to train the deep convolutional network to obtain the final prediction model.

[0089] S5: Input the meteorological forecast data into the photovoltaic prediction neural network model in S4 for photovoltaic prediction. Preferably, the time range of photovoltaic prediction is 1 - 3 days.

[0090] An example of step S1 in the photovoltaic power prediction method of the present invention is as follows:

[0091] Taking the 2014 household photovoltaic data released by the Australian power grid as an example, the historical photovoltaic data of 105 households are selected, the maximum photovoltaic output of each household's photovoltaic in one year is statistically calculated, and it is normalized according to the aforementioned formula, and a typical sunny-day output sequence is selected. Clustering is carried out using the pearson correlation distance as a metric. By adjusting the parameters, some parameters of DBSCAN are finally determined. The neighborhood radius is 0.04, and the minimum number of points is 4. The clustering results are as Figure 2 shown. It can be seen from the bar chart that 105 distributed sites are divided into two clusters, and 9 sites are determined as outliers. The time series diagram of the clustering is as Figure 3 shown. It can be seen from the time series diagram that the output curves can be roughly divided into two strips with a relatively high degree of overlap, and the clustering results basically conform to the senses. The k-nearest neighbor search is used to search for the nearest neighbor sites of 9 abnormal sites, k is taken as 5, and the abnormal sites are assigned to the cluster with the largest number. The final cluster division results are as Figure 4 shown.

[0092] Add the power outputs of each site within the cluster to obtain the total cluster output. Divide the output curve and the corresponding meteorological data by day, with a total of 15 features and 48 time points, forming a 48 * 15 data matrix, and further visualize it, as Figure 5 shown. New data is generated through the constructed data enhancement neural network adversarial learning, as Figure 6 shown. It can be seen that the newly generated data has similar data characteristics to the original data. From the training process, it can be seen that the improved GAN has tended to be stable.

[0093] The original data and the generated data are jointly used to form a training set for training the deep convolutional network. Compare it with the deep convolutional network with the same structure and parameters that only uses the original data. The prediction indicators on the test set are shown in Table 1, and the time series diagram of the prediction results is as Figure 7 shown. It can be seen that this method can effectively improve the prediction accuracy.

[0094] Table 1 Error Statistics of Comparison between Two Methods

[0095]

[0096] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

[0097] The applicant of the present invention has made a detailed description and explanation of the embodiments of the present invention in combination with the accompanying drawings of the specification. However, those skilled in the art should understand that the above embodiments are only the preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, rather than a limitation on the protection scope of the present invention. On the contrary, any improvement or modification made based on the spirit of the present invention should fall within the protection scope of the present invention.

Claims

1. A whole-county photovoltaic prediction method based on cluster division and data augmentation, characterized in that, the whole-county photovoltaic prediction method based on cluster division and data augmentation includes the following steps: S1. Collect the historical sunny-day photovoltaic output data and meteorological data of each photovoltaic output site within the set area and within the set time; S2. Calculate the maximum power of each photovoltaic output site according to the sunny-day photovoltaic output data collected in S1, perform normalization, and then conduct photovoltaic cluster division; S3. Construct a photovoltaic data augmentation neural network model, and the neural network model is an improved GAN; S3.1, set the learning rate α, truncation parameter c, number of batch training samples m, and the number of times the discriminator iterates per iteration of the generator n critic , initialize the discriminator network parameters w t , and the generator network parameters θ t ; S3.2, Check the generator parameter θ t to see if it converges. If it converges, end the iteration; if not, proceed to S3.3; S3.3, Check whether the current iteration count has reached the iteration count threshold n critic , if not, update the discriminator network parameter w t and update the current iteration count, then repeat this step; if the iteration count threshold n is reached critic then proceed to step S3.4; The method for updating the discriminator network parameter w t is as follows: Among them, clip(.) represents the clipping function in deep learning, and RMSProp(.) represents the learning rate adaptive optimizer in deep learning. represents the gradient of the discriminator loss function, and the specific calculation method is as follows: where x (i) is the i-th training batch among the input parameters; z (i) is a batch sample from the sample distribution generated by the generator for the i-th training; is the discriminator network; S3.4, Calculate the generator loss function and update the generator parameters θ t , and return to step 3.2; S4. Perform data augmentation through the photovoltaic data augmentation neural network model constructed in S3 within the photovoltaic cluster, and then input the augmented data into the CNN neural network. After training, a photovoltaic prediction neural network model is obtained; S5: Input the meteorological forecast data into the photovoltaic prediction neural network model in S4 for photovoltaic prediction.

2. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 1, characterized in that, in the S1, the photovoltaic output data is the historical sunny-day photovoltaic output data of the whole county; the meteorological data includes month, day, hour, minute, direct irradiance, diffuse horizontal irradiance, total horizontal irradiance, ambient temperature, air pressure, relative humidity, wind direction, wind speed, surface reflectance, and power generation.

3. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 1 or 2, characterized in that, in the S2, the historical sunny-day photovoltaic output curve is a normalized output curve. First, calculate the maximum power of a single station from the historical output data, and then calculate the normalized value of the curve; The formula for data normalization is as follows: In the above formula, x is the original sample data, and x max is the maximum historical output value of the site, and z is the normalized data.

4. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 3, characterized in that, in the S2, the DBSCAN method is used to divide the photovoltaic clusters. The distance metric of DBSCAN uses the pearson correlation distance; for the abnormal sites generated by DBSCAN during cluster analysis, the k-nearest neighbor algorithm is used to search for the nearest k sites with the pearson correlation distance as the metric, and the cluster numbers of the k sites are counted, and the abnormal sites are divided into the cluster with the largest number of numbers.

5. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 4, characterized in that, The gradient of the generator loss function is: Among them, is the discriminator network, and z (i) is the batch sampling from the sample distribution generated by the generator in the i-th training.

6. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 5, characterized in that, Update the generator parameter θ t The method is as follows: where RMSProp(.) represents the learning rate adaptive optimizer in deep learning.

7. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 1, characterized in that, the data augmentation in S4 means first picturing the data collected in S1, dividing the original data of one day by day, with 48 points per day, forming a picture with a size of 48 * 15 pixels; then performing adversarial learning through the photovoltaic data augmentation neural network model to generate new data with the same distribution.

8. A whole-county photovoltaic prediction method based on cluster division and data augmentation according to claim 1, characterized in that, in the S4, the CNN neural network has a 5-layer structure, and each of the first four layers includes a convolutional layer, a batch normalization layer, and a rectified linear unit layer, and the last layer is a fully connected layer, a batch normalization layer, and a ReLU layer.

Citation Information

Patent Citations

  • Photovoltaic power prediction method

    CN110414748A

  • Photovoltaic power generation power prediction method based on deep belief network

    CN110705760A