Distributed photovoltaic power generation prediction method and device based on partition and time-space correlation
By dividing molecular regions based on AP clustering and deep learning algorithms, a distributed photovoltaic power generation prediction model is constructed, which solves the problem of poor accuracy of regional photovoltaic power prediction and achieves higher precision short-term photovoltaic power prediction.
Patent Information
- Application Number
- CN202510711884.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
The existing regional distributed photovoltaic power prediction has poor accuracy and has failed to effectively use photovoltaic power generation data to mine the spatiotemporal correlation between different sub-regions.
The molecular regions are divided based on the AP clustering algorithm, and the power prediction model is selected to build a power prediction model. The graph convolution network and long and short-term memory network are used to capture spatial and temporal correlation characteristics, and the superscript weight is calculated based on the Pearson coefficient and cosine similarity to perform distributed photovoltaic power generation prediction.
It improves the accuracy and reliability of distributed photovoltaic power generation prediction, can better handle spatiotemporal correlation, and achieve accurate prediction of short-term photovoltaic power.
Smart Images

Figure CN120582089A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photovoltaic power generation prediction, and in particular relates to a distributed photovoltaic power generation prediction method and device based on partitioning and temporal-spatial correlation. Background Art
[0002] As an important form of renewable energy generation, distributed photovoltaic power generation systems have received strong support from national and local governments for their environmental friendliness, flexible layout, and local consumption. However, photovoltaic output suffers from significant volatility, intermittency, and poor power regulation. With the establishment of a large number of distributed photovoltaic power stations, the mismatch between photovoltaic output and the real-time dynamics of existing power generation and consumption has become increasingly prominent, posing a significant challenge to the safe operation of the power grid. Therefore, accurate forecasting of distributed photovoltaic power generation not only provides effective data support for grid energy management, pricing, and load management, but also facilitates the development of reasonable scheduling plans, enabling the effective deployment of distributed photovoltaic power within the power grid, promoting the absorption of large-scale distributed photovoltaic power generation, and improving economic efficiency.
[0003] Currently, various power grid systems require accurate predictions of distributed photovoltaic power generation at different times. Short-term photovoltaic power forecasts, in particular, are widely used in the formulation of daily power generation plans. Numerous methods exist for short-term photovoltaic power forecasting. Most studies utilize satellite cloud images, ground cloud images, and NWP data combined with historical power plant output data. These studies use physical modeling or statistical learning methods to explore correlations between distributed photovoltaic power plants. A variety of prediction models have been developed, including convolutional neural networks, long-short-term memory networks, and graph convolutional neural networks for power forecasting.
[0004] At the same time, because meteorological factors in the same region are roughly the same, the photovoltaic output of distributed photovoltaic power plants may exhibit similar time-varying patterns, which can be defined as spatial similarity. However, due to cloud movement, the photovoltaic output time series corresponding to distributed photovoltaic power plants in different sub-regions may have lag effects or lead effects, a phenomenon defined as temporal correlation.
[0005] Currently, regional distributed photovoltaic power forecasting generally uses indirect scaling methods. These methods require only data from a few power plants within a region to predict the output power of the entire region, and random errors between individual PV plants within the region can offset each other. Most research focuses on accurately predicting power in different subregions using indirect scaling methods after dividing the region. Some studies also focus on the characteristics of prediction errors. However, most studies fail to consider the rational division of subregions with similar output characteristics, or to exploit the spatiotemporal correlations between different subregions using photovoltaic power generation data to perform short-term power forecasting for distributed photovoltaic regions. Summary of the Invention
[0006] The purpose of the present invention is to provide a distributed photovoltaic power generation prediction method and device based on partitioning and temporal and spatial correlation to solve the technical problem of poor accuracy of existing regional distributed photovoltaic power prediction.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation includes the following steps:
[0009] Divide regional distributed photovoltaic power generation into sub-regions based on AP clustering algorithm;
[0010] Based on the results of the sub-region division, representative power stations are selected and power prediction models for the representative power stations under different weather types are constructed;
[0011] The power prediction model of the representative power station under different weather types is used to predict the power of the representative power station, and the predicted power of the representative power station is converted into the predicted power of the sub-region;
[0012] Considering the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the data evaluation score of the sub-region is calculated, and then the superscript weight is determined;
[0013] According to the predicted power of the sub-region and the determined superscript weight, the total predicted power of regional distributed photovoltaic power generation is obtained.
[0014] Furthermore, the steps of dividing the regional distributed photovoltaic power generation into sub-regions based on the AP clustering algorithm are as follows:
[0015] Obtain the original distributed photovoltaic power data and preprocess it to obtain the normalized data of distributed photovoltaic output at each moment;
[0016] Based on the similarity between the output data of distributed power plants and the geographical location information, a similarity matrix is constructed;
[0017] The AP clustering algorithm is used to iteratively calculate the similarity matrix to complete the division of regional distributed photovoltaic sub-regions.
[0018] Furthermore, the distributed photovoltaic output is:
[0019] P(t)=p(t) / Pe
[0020] Where P, p, P e They are the normalized data, real data and rated power of distributed photovoltaic output at each moment respectively.
[0021] Furthermore, the representative power station selection process uses the Pearson coefficient as an output correlation evaluation index to measure the similarity between the outputs of each distributed photovoltaic power station and the corresponding sub-region, and selects the power station with the closest output characteristics as the representative power station.
[0022] Furthermore, the power prediction model is constructed as follows:
[0023] Use graph convolutional network algorithm to mine spatial correlation features of distributed photovoltaic power stations at all times;
[0024] The spatial correlation feature representation given by the graph convolutional network algorithm is input into the long short-term memory network to capture the evolution pattern of spatial correlation features and obtain spatiotemporal correlation features;
[0025] The extracted spatiotemporal correlation features are input into the fully connected layer to generate a power prediction model.
[0026] Furthermore, the steps of calculating the data evaluation score of the sub-region and then determining the superscript weight by considering the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region are as follows:
[0027] The Pearson coefficient and cosine similarity are used to describe the similarity between the output time series of regional power data and sub-regional power data;
[0028] Based on the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the statistical upgrade weight angle is determined and the data evaluation score of the sub-region is calculated;
[0029] The superscript weight is determined based on the data evaluation score of the sub-region.
[0030] Furthermore, the data completeness is described by the number of days with partial data missing and the number of days with full data missing within a day.
[0031] In a second aspect, the present invention provides a distributed photovoltaic power generation prediction system based on partitioning and temporal and spatial correlation, comprising a partitioning module, a construction module, a prediction module, a calculation module and an output module, wherein:
[0032] Division module: used to divide regional distributed photovoltaic power generation into sub-regions based on AP clustering algorithm;
[0033] Construction module: used to select representative power plants based on the sub-region division results and build power prediction models for the representative power plants under different weather conditions;
[0034] Prediction module: used to predict the power of the representative power station using the power prediction model of the representative power station under different weather types, and convert the predicted power of the representative power station into the predicted power of the sub-region;
[0035] Calculation module: used to consider the data integrity, data similarity and distributed photovoltaic capacity ratio of the sub-region, calculate the data evaluation score of the sub-region, and then determine the superscript weight;
[0036] Output module: used to obtain the total predicted power of regional distributed photovoltaic power generation based on the predicted power of the sub-region and the determined superscript weight.
[0037] According to a third aspect, a terminal device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0038] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0039] Compared with the prior art, the present invention has the following beneficial technical effects:
[0040] The present invention provides a distributed photovoltaic power generation prediction method based on zoning and spatiotemporal correlation. It divides regions based on the AP clustering algorithm and constructs a power prediction model for distributed photovoltaic power stations under different weather types. The present invention can better handle distributed photovoltaic systems related to spatiotemporal correlation and improve prediction reliability. It also proposes an evaluation index based on sub-regional data, which can achieve accurate prediction of short-term photovoltaic power.
[0041] Preferably, the original power data is preprocessed by normalization to eliminate the influence of differences in installed capacity of different power stations on clustering.
[0042] Preferably, the weather is divided into three categories: sunny, cloudy and rainy, and the power prediction model is trained separately, which enhances the adaptability of the model to different meteorological conditions.
[0043] Preferably, the power prediction model is built based on GCN-LSTM, which captures the spatial correlation between sub-regions through the graph network structure and captures the temporal evolution using LSTM, and generates high-precision prediction values after fusing the spatiotemporal features, thereby improving the model's ability to express complex spatiotemporal correlations.
[0044] Preferably, the similarity between regional power data and sub-regional power data is introduced as an evaluation index to determine the weight ratio of the statistical scale. By addressing issues such as missing data, the distributed photovoltaic power generation system can be better predicted and managed. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of a distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation in an embodiment of the present invention;
[0046] Figure 2 A detailed flow chart of a distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation in an embodiment of the present invention;
[0047] Figure 3 It is the flow chart of AP clustering algorithm;
[0048] Figure 4 This is the structural flow chart of the GCN-LSTM model. DETAILED DESCRIPTION
[0049] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0050] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0051] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0052] Glossary:
[0053] AP clustering: Affinity Propagation Clustering, is an algorithm for clustering based on the similarity between data points.
[0054] PV: Photovoltaic, usually refers to photovoltaic power generation, is a technology that uses solar energy to convert sunlight into electrical energy.
[0055] GCN: Graph Convolutional Network.
[0056] LSTM: Long Short-Term Memory Network.
[0057] The present invention is described in further detail below with reference to the accompanying drawings:
[0058] like Figure 1 As shown, a distributed photovoltaic power generation prediction method based on partitioning and spatiotemporal correlation includes the following steps:
[0059] Step 1: Divide the regional distributed photovoltaic power generation into sub-regions based on the AP clustering algorithm;
[0060] The specific steps include:
[0061] Obtain the original distributed photovoltaic power data and preprocess it to obtain the normalized data of distributed photovoltaic output at each moment;
[0062] Based on the similarity between the output data of distributed power plants and the geographical location information, a similarity matrix is constructed;
[0063] The AP clustering algorithm is used to iteratively calculate the similarity matrix to complete the division of regional distributed PV (Photovoltaic) sub-regions.
[0064] Among them, the original distributed photovoltaic power data is preprocessed according to the difference in installed capacity;
[0065] The pre-processed photovoltaic output is:
[0066] P(t)=p(t) / P e (1)
[0067] Where P, p, P e The normalized data, actual data, and rated power of distributed photovoltaic output at each moment are shown in Figure 2. When the rated power information of a distributed photovoltaic power station is lacking, the maximum output power of the power station can be used instead of the rated power.
[0068] The Euclidean distance matrix H between the output data of photovoltaic power plants is constructed by the photovoltaic power data sample matrix P, and the similarity matrix S is constructed using the negative of the square of the Euclidean distance matrix H.
[0069] The attraction matrix r and attribution matrix a are calculated by the similarity matrix. The calculation formula is:
[0070] r(i,x)=s(i,x)-max[a(i,t)+s(i,t)](t≠x) (5)
[0071]
[0072] r(i,x) describes the suitability of data point x as the cluster center of point i, taking into account other possible cluster centers of point i. a(i,x) reflects the suitability of point i to choose x as its cluster center, taking into account the support of point x as a cluster center other than point i.
[0073] like Figure 3 As shown, in step S104, the steps of the AP algorithm are:
[0074] 1) Establish a sample matrix based on the data set, calculate the similarity matrix S, and set the median of the similarity matrix as the reference degree Z;
[0075] 2) Initialize the attribution matrix a(i,k) = 0 and calculate the gravity matrix r according to formula (5);
[0076] 3) Using the attraction matrix r, calculate the attribution matrix a according to formula (6);
[0077] 4) Iteratively update the r and a matrices and determine the algorithm's convergence based on the termination criteria. The sample point's attraction and attribution information are summed to test the sample point's decision to select the cluster center. If the cluster center remains unchanged after multiple iterations, or if the number of iterations exceeds the set number, the algorithm terminates.
[0078] In step S201, representative power plants are selected based on the principle of maximizing the Pearson correlation coefficient. The calculation formula for similarity evaluation is:
[0079]
[0080] Among them, ρ(P i ,P j ) is the correlation coefficient between the two time series, ranging from [-1, 1]; r is the length of the time series; P i,t (t=t0,t 0+1 ,…,n) is the PV power output time series of the photovoltaic power station, and its mean is P i ;P j,t (t=t0,t 0+1 ,…,n) is the PV power output time series of the sub-region, and its mean is P j ; The power plant with the largest correlation coefficient with the sub-region output is selected as the representative power plant of each sub-region. The larger the correlation coefficient, the stronger the correlation.
[0081] Step 2: Based on the results of the sub-region division, select representative power stations and build power prediction models for the representative power stations under different weather types;
[0082] The detailed steps are:
[0083] The Pearson coefficient is used as the output correlation evaluation index to measure the similarity between the output of each distributed photovoltaic power station and the corresponding sub-region, and the power station with the closest output characteristics is selected as the representative power station;
[0084] Classify weather types into three categories: sunny, cloudy, and rainy. Then divide the dataset into three categories. Use the data from each category to train a photovoltaic power generation prediction model for the corresponding power station under the corresponding weather type.
[0085] The graph convolutional network (GCN) algorithm is used to mine the spatial correlation features of distributed photovoltaic power stations at each moment. The spatial correlation feature representation given by the GCN algorithm is then input into the long short-term memory (LSTM) network to capture the evolution pattern of the spatial correlation features and obtain the spatiotemporal correlation features.
[0086] The extracted spatiotemporal correlation features are input into the fully connected layer to generate reliable and high-quality prediction results to predict the power generation of representative power plants.
[0087] Step 3: Use the power prediction model of the representative power station under different weather types to predict the power of the representative power station, and convert the predicted power of the representative power station into the predicted power of the sub-region;
[0088] Before converting the predicted power of the representative power station into the predicted power of the sub-region, the conversion factor of the rated capacity of the representative power station and the sub-region is calculated.
[0089] Step 4: Considering the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the data evaluation score of the sub-region is calculated, and then the superscript weight is determined;
[0090] Correspondingly, the Pearson coefficient and cosine similarity are used to describe the similarity between the output time series of regional power data and sub-regional power data;
[0091] Based on the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the statistical upgrade weight angle is determined and the data evaluation score of the sub-region is calculated;
[0092] The superscript weight is determined based on the data evaluation score of the sub-region.
[0093] Step 5: According to the predicted power of the sub-region and the determined superscript weight, the total predicted power of regional distributed photovoltaic power generation is obtained.
[0094] In another embodiment of the present invention, a distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation is provided. Figure 2 As shown, the following steps are included:
[0095] S1: By analyzing the spatiotemporal correlation between the output time series of distributed photovoltaic power stations in sub-regions, a sub-region division method based on AP clustering is proposed to provide support for power prediction by mining the spatiotemporal correlation information between sub-regions;
[0096] S2: Based on the regional division results, representative power plants are selected, and the weather type division and the spatiotemporal correlation of output between subregions are considered. Using the historical power data and weather type of the forecast day as input parameters, a graph network structure is established to mine its spatiotemporal correlation characteristics, thereby obtaining a power prediction model for distributed PV power plants under different weather types.
[0097] S3: A regional distributed photovoltaic statistical forecasting method based on sub-regional data evaluation is proposed. The sub-regional data are scored using the Pearson coefficient, cosine similarity, number of missing output data, and distributed photovoltaic capacity, and the superscript weights are determined to improve the accuracy of regional power forecasting.
[0098] The detailed steps of step S1 include:
[0099] S101, obtaining original distributed photovoltaic power data;
[0100] S102. Preprocessing the original distributed photovoltaic power data based on the installed capacity difference to obtain normalized data of the distributed photovoltaic output at each moment;
[0101] S103, constructing a similarity matrix using Euclidean metric to describe the similarity between the output of each distributed power plant and its geographical location;
[0102] S104: Use the AP clustering algorithm to complete the sub-region division of the regional distributed PV.
[0103] like Figure 4 As shown, the detailed steps of step S2 include:
[0104] S201. Use the Pearson coefficient as an output correlation evaluation indicator to measure the similarity between the outputs of each distributed photovoltaic power station and the corresponding sub-region, and select the power station with the closest output characteristics as the representative power station;
[0105] S202, classifying weather types into three types: sunny, cloudy, and rainy, and then dividing the data set into three categories, using the data of each type to train a photovoltaic power generation prediction model representing the power station under the corresponding weather type;
[0106] S203, using a graph convolutional network (GCN) algorithm to mine spatially relevant features of the distributed photovoltaic power station at each moment; then inputting the feature representation given by the GCN algorithm into a long short-term memory (LSTM) network to capture the evolution pattern of the spatially relevant features;
[0107] S204: Input the extracted spatiotemporal correlation features into the fully connected layer to generate reliable high-quality prediction results and predict the power generation of the representative power plant.
[0108] The detailed steps of step S3 include:
[0109] Through regional data evaluation indicators, regional data are evaluated from three aspects: data integrity, data similarity and distributed photovoltaic capacity;
[0110] S301, introducing the similarity between regional power data and sub-regional power data as one aspect of the evaluation index to determine the weight ratio of the statistical scale, and using the Pearson coefficient and cosine similarity to describe the similarity between the two PV output time series;
[0111] S302: Considering the data integrity, data similarity and distributed photovoltaic capacity ratio of the sub-region, calculate the data evaluation score S of the sub-region from the perspective of determining a more appropriate statistical upgrade weight. j ;
[0112] S303: Convert the predicted power of the representative power station into the predicted power of the sub-region, and then use the data evaluation score of each sub-region to determine the weight ratio on the statistical scale, so as to obtain the total predicted power of regional distributed photovoltaic power generation.
[0113] In step S102, the pre-processed photovoltaic output is:
[0114] P(t)=p(t) / P e (1)
[0115] Where P, p, P e The normalized data, actual data, and rated power of distributed photovoltaic output at each moment are shown in Figure 2. When the rated power information of a distributed photovoltaic power station is lacking, the maximum output power of the power station can be used instead of the rated power.
[0116] In step S103 , the Euclidean distance matrix H between the output data of the photovoltaic power station is constructed by the photovoltaic power data sample matrix P, and the similarity matrix S is constructed by using the negative of the square of the Euclidean distance matrix H.
[0117] The sample matrix P is expressed as:
[0118]
[0119] The Euclidean distance matrix H is expressed as:
[0120]
[0121] Where u, w=1, 2, ..., n, d(P u ,Pw ) represents the Euclidean distance between the PV output data of power plants u and w, d(L u ,L w ) is the Euclidean distance calculated from the latitude and longitude of the two power plants. K cdi is the weight coefficient, which is used to adjust the weight ratio of geographic location similarity.
[0122] The similarity matrix S is expressed as:
[0123]
[0124] The attraction matrix r and attribution matrix a are calculated by the similarity matrix. The calculation formula is:
[0125] r(i,x)=s(i,x)-max[a(i,t)+s(i,t)](t≠x) (5)
[0126]
[0127] r(i,x) describes the suitability of data point x as the cluster center of point i, taking into account other possible cluster centers of point i. a(i,x) reflects the suitability of point i to choose x as its cluster center, taking into account the support of point x as a cluster center other than point i.
[0128] like Figure 3 As shown, in step S104, the steps of the AP algorithm are:
[0129] 1) Establish a sample matrix based on the data set, calculate the similarity matrix S, and set the median Z of the similarity matrix as the reference degree;
[0130] 2) Initialize the attribution matrix a(i,k) = 0 and calculate the gravity matrix r according to formula (5);
[0131] 3) Using the attraction matrix r, calculate the attribution matrix a according to formula (6);
[0132] 4) Iteratively update the r and a matrices and determine the algorithm's convergence based on the termination criteria. The sample point's attraction and attribution information are summed to test the sample point's decision to select the cluster center. If the cluster center remains unchanged after multiple iterations, or if the number of iterations exceeds the set number, the algorithm terminates.
[0133] In step S201, representative power plants are selected based on the principle of maximizing the Pearson correlation coefficient. The calculation formula for similarity evaluation is:
[0134]
[0135] Among them, ρ(P i ,Pj ) is the correlation coefficient between the two time series, ranging from [-1, 1]; r is the length of the time series; P i,t (t=t0,t 0+1 ,…,n) is the PV power output time series of the photovoltaic power station, and its mean is P i ;P j,t (t=t0,t 0+1 ,…,n) is the PV power output time series of the sub-region, and its mean is P j ; The power plant with the largest correlation coefficient with the sub-region output is selected as the representative power plant of each sub-region. The larger the correlation coefficient, the stronger the correlation.
[0136] In step S202, a deep learning algorithm is used to calculate the contribution of climate factors to photovoltaic power, and weather features with a strong correlation with photovoltaic power are screened out. The weather data is divided into three types of weather: sunny, cloudy and rainy days through cluster analysis.
[0137] In step S203, a representative power plant power prediction model is calculated based on the GCN algorithm and the LSTM algorithm, using the r-step historical PV power output data p i,n , then the predicted q-step PV power output data p i,t+q The prediction model can be expressed as:
[0138]
[0139] p i,r =[p i,t-r+1 ,p i,t-r+2 ,…,p i,t ] (8)
[0140] The f function is a proposed power prediction model based on the GCN and LSTM algorithms. The essence of GCN is to extend convolution to non-Euclidean data. By taking the derivative of the complex Laplace matrix, we can derive the expression for graph convolution. Then, in the GCN-LSTM model, the learned spatially correlated features are sent to the LSTM layer. The LSTM layer has a strong ability to learn long-term dependencies in sequential data and can capture the evolutionary patterns of weighted dynamic networks. A standard LSTM architecture can be described as a structure encapsulating multiple multiplication gate units. For a given time step t, the LSTM unit takes as input the current input vector and the state vector of the previous time step, and then outputs the state vector for the current time step.
[0141] In step S301, the integrity of distributed photovoltaic output data is described by the number of days with partial data missing and the number of days with full data missing. The number of days with partial data missing is recorded as S1, and the number of days with full data missing is recorded as S2. At the same time, the Pearson coefficient and cosine similarity are used to describe the similarity between two PV output time series. The Pearson coefficient is recorded as S3, and the cosine similarity S4 is expressed as follows:
[0142]
[0143] Where p j,t (t=t0,t 0+1 ,…,n) is the time series of regional distributed photovoltaic power generation output; p all,t (t=t0,t 0+1 ,…,n) is the power output time series of regional distributed photovoltaic; r is the length of the time series; S4(P j ,P all ) is the cosine similarity of the two time series, between -1 and 1. A cosine similarity equal to -1 means that the two vectors point in opposite directions, that is, the two time series are completely dissimilar; a cosine similarity equal to 1 means that their points are exactly the same and the two time series are completely similar.
[0144] In step S302, the data evaluation score S j The formula is:
[0145]
[0146] Where, P Ni is the rated capacity of each distributed photovoltaic power station in the sub-region, i = 1, 2, ..., M, M is the number of distributed photovoltaic power stations in the sub-region; P Nj is the sum of the rated capacities of all distributed photovoltaics in sub-region j, j = 1, 2, ..., L, L is the number of sub-regions in the region, and a is the evaluation coefficient, which is used to adjust the score range to between -1 and 1.
[0147] In step S303, the conversion coefficients representing the rated capacities of power plants and sub-regions are calculated, and the scale prediction capabilities of the sub-regions are calculated. Finally, the evaluation scores of each sub-region are used to calculate the statistical scale weights.
[0148] Calculate the conversion factor representing the rated capacity of the power station and sub-region using the following formula:
[0149]
[0150] The calculation formula of regional scale prediction ability is:
[0151]
[0152] Where, is the prediction capability of the power station in sub-region j, is the prediction ability of sub-region j.
[0153] Finally, the evaluation scores of each sub-region are used to calculate the statistical weight of the scale. The calculation formula is:
[0154]
[0155] Where θ j is the statistical upscaling weight of sub-region j, is the total predictive power for the region.
[0156] In another embodiment of the present invention, a distributed photovoltaic power generation prediction system based on partitioning and temporal and spatial correlation is provided, comprising a partitioning module, a construction module, a prediction module, a calculation module, and an output module, wherein:
[0157] Division module: used to divide regional distributed photovoltaic power generation into sub-regions based on AP clustering algorithm;
[0158] Construction module: used to select representative power plants based on the sub-region division results and build power prediction models for the representative power plants under different weather conditions;
[0159] Prediction module: used to predict the power of the representative power station using the power prediction model of the representative power station under different weather types, and convert the predicted power of the representative power station into the predicted power of the sub-region;
[0160] Calculation module: used to consider the data integrity, data similarity and distributed photovoltaic capacity ratio of the sub-region, calculate the data evaluation score of the sub-region, and then determine the superscript weight;
[0161] Output module: used to obtain the total predicted power of regional distributed photovoltaic power generation based on the predicted power of the sub-region and the determined superscript weight.
[0162] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0163] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0164] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that after reading the present invention, those skilled in the art may still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the scope of protection of the pending claims of the invention.
Claims
1. A distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation is characterized by: The following steps are involved: Divide regional distributed photovoltaic power generation into sub-regions based on AP clustering algorithm; Based on the results of the sub-region division, representative power stations are selected and power prediction models for the representative power stations under different weather types are constructed; The power prediction model of the representative power station under different weather types is used to predict the power of the representative power station, and the predicted power of the representative power station is converted into the predicted power of the sub-region; Considering the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the data evaluation score of the sub-region is calculated, and then the superscript weight is determined; According to the predicted power of the sub-region and the determined superscript weight, the total predicted power of regional distributed photovoltaic power generation is obtained.
2. A distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 1, characterized in that: The steps of dividing regional distributed photovoltaic power generation into sub-regions based on the AP clustering algorithm are as follows: Obtain the original distributed photovoltaic power data and preprocess it to obtain the normalized data of distributed photovoltaic output at each moment; Based on the similarity between the output data of distributed power plants and the geographical location information, a similarity matrix is constructed; The AP clustering algorithm is used to iteratively calculate the similarity matrix to complete the division of regional distributed photovoltaic sub-regions.
3. A distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 2, characterized in that: The distributed photovoltaic output is: P(t)=p(t) / P e Where P, p, P e They are the normalized data, real data and rated power of distributed photovoltaic output at each moment respectively.
4. The distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 1 is characterized in that: The representative power station selection process uses the Pearson coefficient as an output correlation evaluation index to measure the similarity between the outputs of each distributed photovoltaic power station and the corresponding sub-region, and selects the power station with the closest output characteristics as the representative power station.
5. The distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 1 is characterized in that: The construction process of the power prediction model is as follows: Use graph convolutional network algorithm to mine spatial correlation features of distributed photovoltaic power stations at all times; The spatial correlation feature representation given by the graph convolutional network algorithm is input into the long short-term memory network to capture the evolution pattern of spatial correlation features and obtain spatiotemporal correlation features; The extracted spatiotemporal correlation features are input into the fully connected layer to generate a power prediction model.
6. The distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 1 is characterized in that: The steps of calculating the data evaluation score of the sub-region and then determining the superscript weight by considering the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region are as follows: The Pearson coefficient and cosine similarity are used to describe the similarity between the output time series of regional power data and sub-regional power data; Based on the data integrity, data similarity and distributed photovoltaic capacity proportion of the sub-region, the statistical upgrade weight angle is determined and the data evaluation score of the sub-region is calculated; The superscript weight is determined based on the data evaluation score of the sub-region.
7. A distributed photovoltaic power generation prediction method based on partitioning and temporal and spatial correlation according to claim 6, characterized in that: The data completeness was described by the number of days with partial missing data and the number of days with full missing data.
8. A distributed photovoltaic power generation prediction system based on partitioning and temporal and spatial correlation, characterized in that: The distributed photovoltaic power generation prediction method based on partitioning and spatiotemporal correlation according to any one of claims 1 to 7 comprises a partitioning module, a construction module, a prediction module, a calculation module and an output module, wherein: Division module: used to divide regional distributed photovoltaic power generation into sub-regions based on AP clustering algorithm; Construction module: used to select representative power plants based on the sub-region division results and build power prediction models for the representative power plants under different weather conditions; Prediction module: used to predict the power of the representative power station using the power prediction model of the representative power station under different weather types, and convert the predicted power of the representative power station into the predicted power of the sub-region; Calculation module: used to consider the data integrity, data similarity and distributed photovoltaic capacity ratio of the sub-region, calculate the data evaluation score of the sub-region, and then determine the superscript weight; Output module: used to obtain the total predicted power of regional distributed photovoltaic power generation based on the predicted power of the sub-region and the determined superscript weight.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.