A Method and System for Generating Distributed Photovoltaic Random Output Scenarios
Through graph convolution generation adversarial networks and clustering algorithms, the photovoltaic output scenarios are generated and reduced, which solves the problem of difficult to capture the spatial and temporal correlation of photovoltaic units, and achieves more accurate and diversified photovoltaic output scenario generation.
Patent Information
- Application Number
- CN202411866902.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-12-18
AI Technical Summary
When the existing photovoltaic output scenario generation method deals with power systems in the context of high proportion of renewable energy, it is difficult to effectively capture the spatial and temporal correlation between photovoltaic units, resulting in large deviations from the actual situation and difficulty in meeting the diversity requirements.
Graph convolution generation adversarial network (GCGAN) is used to combine ST-DBSCAN and K-medoids clustering algorithms to implicitly model the spatial and temporal characteristics of photovoltaic sites, generate and reduce photovoltaic output scenarios, and preserve the spatiotemporal correlation of photovoltaic units.
A more complete and comprehensive photovoltaic output scenario was generated, which improved the accuracy of scene generation, and preserved the spatial and temporal correlation of photovoltaic sites through the reduction method of space and time dimensions, and reduced the calculation amount.
Smart Images

Figure CN119988857B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photovoltaic scenario generation, and in particular to a method and system for generating distributed photovoltaic random output scenarios. Background Art
[0002] With the continuous increase in the penetration rate of renewable energy, especially large-scale integration, the uncertainty of its power generation has brought huge challenges to the operation and planning of power systems. The scenario generation method is an effective method for dealing with the uncertainty of renewable energy power generation and is widely used in stochastic optimization problems such as power system economic dispatch, energy storage configuration, and market bidding. The basic idea of the scenario generation method is to transform various uncertain factors into a combination of multiple deterministic factors, and each combination represents a scenario. Then, each deterministic scenario is optimized to reduce the difficulty of modeling and solving. The key of the scenario generation method lies in how to select a representative set of scenarios to ensure the accuracy of the model.
[0003] The scenario generation method includes scenario generation and scenario reduction. Regarding the research on scenario generation, numerous methods have been proposed in the academic community. The most commonly used is the probability model-based method, which analyzes historical data to establish a probability distribution that conforms to the characteristics of stochastic power output, and generates scenarios through sampling. Pinson P, based on the probabilistic prediction of wind power, uses Gaussian Copula to fit the interdependent structure between wind power prediction errors and prediction distributions. Regarding the scenario analysis research of multiple power generation locations, Li Jinghua considered the spatial correlation of power generation in multiple wind farms, avoided the direct construction of the joint distribution function of multiple wind farms, and proposed a Copula-based joint scenario simulation method to effectively characterize the power generation scenarios of multiple geographical locations; Scholar Jin Tianran established a time correlation model between wind power and photovoltaic based on multivariate normal distribution, and at the same time established a Copula model of vine structure, and used the vine structure of multi-dimensional variables for sampling to generate scenarios. The existing multi-site joint output based on model probability often generates the output of a single photovoltaic site first, and then uses the copula function to construct the joint probability function between multiple sites. However, the scenario analysis method based on the probability model often needs to assume the law characteristics of unknown distribution data according to statistical experience, and requires artificial selection of model functions. The rationality of probability assumptions is related to the quality and reliability of scenario analysis, and the generated scenarios often deviate greatly from the actual scenarios and cannot meet the requirements of generating scenario diversity.
[0004] Another popular method is the time series method. Zou Bin used the Auto Regressive Moving Average (ARMA) model to characterize wind speed and simulate stochastic production based on the effective capacity distribution of generator sets. Lei Yu used the ARMA model to model the dynamic power output of wind power to generate an initial scenario set. In the scenario simulation study of multiple power generation locations, Duong D. Le used the time series method to generate scenarios for multiple spatially correlated locations, enabling the generated scenarios to maintain the temporal correlation of the time series and characterizing the spatial correlation of different geographical locations. However, autoregressive models and state space representations are prone to problems such as overfitting of the model and pattern recognition errors, and it is difficult to use time series models to characterize the diversity of renewable energy power generation dynamics.
[0005] In recent years, machine learning methods have also begun to be applied to scenario generation. Vagropoulos S I proposed using an Artificial Neural Network (ANN) to generate scenarios and applied it to the scenario generation of power load, solar energy, and wind power. Cui Mingjian used a neural network to generate scenarios and used the generated scenarios to predict the ramping events of wind power. With the development of deep learning, more and more deep learning-based generative models are being used in scenario generation. Yize Chen proposed a renewable energy scenario generation method based on a generative adversarial network, which can correctly capture the temporal and spatial correlations of solar and wind power generation and can efficiently generate data that conforms to the observational characteristics by adding conditional label information. Dong Xiaochong proposed a day-ahead scenario generation method based on a conditional generative adversarial network to characterize the uncertainty of the day-ahead output of renewable energy by satisfying the relationship between the noise distribution that meets the prediction conditions and the prediction scenario. Generative models provide a new solution for capturing the true distribution of data. Without analyzing the distribution characteristics of the data and without probability assumptions, they can learn the unknown distribution of historical data in an unsupervised manner and generate data samples with the same distribution characteristics as the training data.
[0006] Regarding the research on scenario reduction for photovoltaic power plants and wind farms, Wang Qun, Yang Xian, etc. respectively used improved K-medoids clustering and improved K-means clustering algorithms for scenario reduction to improve the speed and accuracy of scenario reduction. J. Dupaová proposed the predecessor and successor algorithm based on probability distance, and Heitsch H improved the algorithm on the basis of J. Dupaová and proposed the fast predecessor elimination algorithm for scenario reduction to improve the calculation speed and accuracy. Heitsch H also proposed the synchronous back substitution reduction method, which is now widely used in scenario reduction. However, existing photovoltaic output scenario reduction often only performs reduction in the time series dimension or the spatial dimension alone, ignoring the spatio-temporal correlation between photovoltaic units. Summary of the Invention
[0007] The object of the present invention is to provide a method for generating distributed photovoltaic random output scenarios. Aiming at studying power system problems under the background of high proportion of renewable energy, a large number of original photovoltaic output scenarios are reduced to a few representative scenarios to achieve the purpose of streamlining data and reducing subsequent calculation amounts, transforming the uncertainty problem into a deterministic problem that can be effectively solved, and providing a data basis for subsequent power system planning, operation, dispatching and other problems.
[0008] To achieve the above object, the present invention provides the following technical solution: A method for generating distributed photovoltaic random output scenarios, comprising the following steps:
[0009] S1. Collect historical output data of each photovoltaic site, and perform normalization and standardization processing on the historical output data;
[0010] S2. Use noise as the input of the generator in the graph convolutional generative adversarial network, use the generator to generate a large amount of output data, use the output data generated by the generator and the processed historical output data as the input of the discriminator in the graph convolutional generative adversarial network, and train the graph convolutional generative adversarial network until the discriminator cannot determine whether the data comes from the generator or the real historical output data, and the generative adversarial network reaches Nash equilibrium, and it is considered that the samples generated by the generator are basically the same as the real data in statistical distribution at this time;
[0011] The noise is data sampled from a simple distribution, and the simple distribution includes Gaussian distribution or uniform distribution;
[0012] S3. Input the noise into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios, and complete the data augmentation of the output scenarios;
[0013] S4. Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish the data with concentrated density and the scattered data, select typical photovoltaic sites, and splice the outputs of each typical photovoltaic site according to the time dimension;
[0014] S5. Use the K-medoids clustering method to select typical output dates, and select the generated output scenarios from the completely generated photovoltaic output scenarios according to the dates to obtain typical photovoltaic output scenarios.
[0015] In some embodiments, in S1, the historical output data of the photovoltaic site is divided into the rainy season and the dry season, the data format is set as a three-dimensional matrix, and the photovoltaic power information of all sites every day is recorded in the form of a two-dimensional matrix {x n,t}.
[0016] In some embodiments, the layers in the generator and discriminator of the graph convolutional generative adversarial network all use graph convolutional networks. The generator and discriminator both include graph filters in the spatial dimension and feature filters in the temporal dimension. In the spatial dimension, graph filters are used to capture the spatial correlation relationships between photovoltaic sites; in the temporal dimension, one-dimensional convolutional filters are used to process the temporal features to complete feature extraction.
[0017] In some embodiments, the following steps are included:
[0018] Use the input data matrix as the parameter of the feature filter:
[0019]
[0020] The last layer uses the sigmoid function to generate an output value in the range of (0, 1) as the photovoltaic output data;
[0021] Through the feature aggregation of adjacent nodes by graph convolution, use the exponentially transformed correlation coefficient as the weight of the graph filter. The calculation expression is:
[0022]
[0023] Regard the matrix multiplication term in formula (6) as the output of the feature filter, and use a one-dimensional convolutional filter with a length of and trainable weight coefficients. The element in the j-th column is:
[0024]
[0025] In the formula, is the input data matrix with features for each layer, where L is the total number of hidden layers in the generator; σ(·) is the non-linear activation function; A is the graph filter, is the weight matrix; C i,j is the correlation coefficient between node i and node j; m is the position offset of the sliding window in the one-dimensional convolution; is half of the width of the convolution window; is the position feature matrix for the m offset; is the feature of the j-th column of the input matrix at the m offset position.
[0026] In some embodiments, the loss function of the generator is:
[0027]
[0028] The loss function of the discriminator is:
[0029]
[0030] where z is a noise variable; P z is a simple distribution; G(z) is the mapping of the noise variable to a generated data space; P r is the true sample probability distribution; x is the generated sample; G is the generator; D is the discriminator; P G is the generated sample probability distribution; E is the expectation.
[0031] In some embodiments, in S4, selecting typical photovoltaic sites includes the following steps:
[0032] S41. Set the input domain radius Eps, the minimum number minpts, and the time neighborhood Δt;
[0033] S42. Starting from an arbitrary unlabeled point, select a point p as the core point for scanning. Scan the neighborhood within the time neighborhood Δt of the core point p, and count the number of neighborhood points that satisfy the distance threshold in the spherical region Eps(s) with the core point p as the center and the distance neighborhood as the radius within the scanned time neighborhood Δt;
[0034] S43. If the number of neighborhood points of the core point p is less than the set threshold MinPts, mark this point as a noise point;
[0035] S44. If the number of neighborhood points is greater than the set threshold MinPts, mark this point as the core point p, generate a cluster number as cluster C1, and add all the points of the core point p in the neighborhood Eps(s) to the cluster C1 to be scanned;
[0036] S45. For the unlabeled points within cluster C1, repeat steps S42 - S44 until all the points within cluster C1 are marked;
[0037] S46. Continue to select marked points from the data set, repeat steps S42 - S45, and cluster them into another class until all points are marked.
[0038] In some embodiments, the specific way of splicing the output of each typical photovoltaic site according to the time dimension is: splicing the typical photovoltaic sites according to the corresponding dates to form a one - dimensional sequence:
[0039] V d =[P1(1,d),P1(2,d),...,P1(t,d),...,P i (1,D),P i (2,D),...,P i (t,D)] T (9);
[0040]
[0041] In the formula, P i (t, D) is the output of the i-th typical site at the t-th hour on the D-th day; t = 1, 2,..., 24, d = 1, 2,..., D, D is the total number of days, D = 365; i = 1, 2,..., M, where M is the number of typical sites; is a 24×M matrix.
[0042] In some embodiments, S5 includes the following steps:
[0043] S51. Randomly select r sequences from all sequences as the initial clustering centers;
[0044] S52. According to the principle of the closest distance to the clustering center, assign the remaining objects to each class, and calculate the distance between the two closest sequences;
[0045] S53. Minimize the total weighted distance from all scenarios to the clustering center;
[0046] S54. Determine whether it converges. If it does not converge, repeat S52. If it has converged, then the r clustering centers obtained by clustering are the typical days of photovoltaic output;
[0047] S55. Calculate the probability of each sequence appearing after clustering.
[0048] In some embodiments, the calculation expression for calculating the distance between the two closest sequences is:
[0049]
[0050] In the formula, d(u i , u j ) is the distance between two sequences. For each sequence u i , select the clustering center u j with the smallest distance to it to calculate the distance; T is the number of time points of the sequence; is the value of scenario u i at the t-th time point; the distance d(u i , u j ) is the sum of the absolute values of the differences between the two sequences at each time point.
[0051] A distributed photovoltaic random output scenario generation system, applying the distributed photovoltaic random output scenario generation method as described above, includes:
[0052] Data acquisition module: used to acquire the historical output data of each photovoltaic site, and perform normalization and standardization processing on the historical output data;
[0053] Graph Convolutional Generative Adversarial Network Processing Module: Use noise as the input of the generator, and use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the Graph Convolutional Generative Adversarial Network Processing Module to train the Graph Convolutional Generative Adversarial Network;
[0054] Input the noise into the generator of the trained Graph Convolutional Generative Adversarial Network Processing Module to generate a photovoltaic output scenario, completing the data augmentation of the output scenario;
[0055] Spatial Clustering Processing Module: Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish the data with concentrated density and the scattered data, select typical photovoltaic sites, and splice the outputs of each typical photovoltaic site according to the time dimension;
[0056] Temporal Clustering Processing Module: Use the K-medoids clustering method to select typical output dates, and select the generated output scenarios from the completely generated photovoltaic output scenarios according to the dates to obtain typical photovoltaic output scenarios.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] 1. Based on the Graph Convolutional Generative Adversarial Network, the present invention uses the method of implicit modeling to mine the high-dimensional non-linear features of historical data. It uses graph filters in the spatial dimension and one-dimensional convolutional filters in the temporal dimension to complete feature extraction, learning the time-space correlation and meteorological correlation contained in the outputs of multiple photovoltaic units, and improving the accuracy of generating uncertain photovoltaic output scenarios. At the same time, the Graph Convolutional Generative Adversarial Network learns the existing photovoltaic output data distribution, enabling the generator to learn to simulate various potential output situations, including marginal and rare events; the discriminator promotes the generator to improve the quality, making the generated results closer and closer to the real output distribution, thus generating a more complete and comprehensive photovoltaic output scenario and completing the data augmentation of the output scenario.
[0059] 2. The present invention proposes a method for reducing the output scenarios of multiple photovoltaic sites. It uses spatial clustering to obtain typical sites, reconstructs the output data of typical sites, and then clusters the time to obtain typical output dates. The complete output of each site corresponding to the typical output dates is the typical output scenario. This method reduces the scenarios from two dimensions of space and time, retaining the spatial and temporal correlations of the outputs of each photovoltaic site and obtaining typical output scenarios with the output data of multiple photovoltaic sites. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0061] Figure 2Schematic diagram of the data form of the input generation adversarial network in Embodiment 1 of the present invention;
[0062] Figure 3 Schematic diagram of the graph convolutional generative adversarial network structure in Embodiment 1 of the present invention;
[0063] Figure 4 Schematic diagram of the data reconstruction form in one embodiment of the present invention;
[0064] Figure 5 Schematic diagram of the system structure in Embodiment 2 of the present invention. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0066] Embodiment 1
[0067] Please refer to Figures 1-4 , a method for generating a distributed photovoltaic random output scenario, including the following steps:
[0068] I. Generation of photovoltaic output scenarios based on graph convolutional generative adversarial networks.
[0069] (1). Data processing:
[0070] In this embodiment, first, the historical output data of each photovoltaic site is collected, and the historical photovoltaic data is divided into the rainy season and the dry season.
[0071] The input data for generating the photovoltaic output curve scenario is the photovoltaic output power information of N sites in one year. There are 365 days in a year. The data format is converted into a three-dimensional matrix, and the data size is expressed as N×T×D, where N is the total number of photovoltaic sites; T takes 24, representing 24 hours of data; D takes 365, representing 365 days in a year. The photovoltaic power information of all sites on each day can be recorded in the form of a two-dimensional matrix {x n,t}, where n is the number of sites, t represents the time, and x n,t represents the photovoltaic output power information of the nth site at time t. The input matrix form is shown as:
[0072]
[0073] Before inputting data into the model, it is also necessary to perform feature scaling on the data, map the original data to a specific range, remove units and unify the dimensions, and turn it into pure data samples, that is, normalize and standardize the data. Data normalization is to limit the data within a certain range, usually mapping the data to the range of -1 to 1 or 0 to 1. The former is called Min-Max normalization, and the commonly used Min-Max normalization calculation expression is:
[0074]
[0075] The data obtained through the above operations is in the form of Figure 1 as shown.
[0076] (2). Use the Graph Convolutional Generative Adversarial Network (GCGAN) to generate photovoltaic output scenarios.
[0077] The generative adversarial network is a deep learning framework, including a generative network and a discriminative network.
[0078] The goal of the generative adversarial network is to generate real samples so that the discriminative network has difficulty distinguishing its authenticity. To achieve this goal, it is necessary to define the loss function for training the generative network and the loss function for training the discriminative network.
[0079] When the discriminative network is fixed and the generative network is trained, the probability of the generated sample G(z) should be greater, and at the same time the loss function L G is smaller.
[0080] The training goal of the discriminative network is to distinguish P r and P G . When the parameters θ G of the generative network are fixed and the parameters θ D of the discriminative network are updated, the greater the discrimination difference indicates the stronger the judgment ability of the discriminative network, that is, to maximize the difference between E[D(·)] and E[D(G(·))]. According to the basic principle of the generative adversarial network, when the discriminative network can well distinguish the input samples, its loss function is smaller.
[0081] The loss function of the generator is:
[0082]
[0083] The loss function of the discriminator is:
[0084]
[0085] In the formula, z is a noise variable; P zis a simple distribution; G(z) maps a noise variable to a generated data space; P r is the probability distribution of real samples; x is a generated sample; G is a generator; D is a discriminator; P G is the probability distribution of generated samples; E is the expectation; E[D(·)] is the expectation of the discriminator for real data; E[D(G(·))] is the expectation of the discriminator for accepting generated data.
[0086] The zero-sum game participated by the two networks is denoted as V(D, G). For the optimization of the loss function, it is necessary to update according to the update of the other network while only updating the parameters of its own network. The two networks are continuously trained alternately and their parameters are updated continuously until the adversarial networks reach Nash equilibrium. When the training converges, a pair of optimal parameters can be obtained. Because there are multiple extreme points in the high-dimensional data space, the parameters at this time and are two extreme points of V(D, G). The principle of the generative adversarial network is a minimax optimization problem, defined by the following formula:
[0087]
[0088] In the formula, V(D, G) is a binary cross-entropy function, and the ultimate goal of this function is to minimize the JS distance between the probability distribution P G of the generated samples and the probability distribution P r of the real samples.
[0089] As Figure 2 shown, the layers in the generator and discriminator of the graph convolutional generative adversarial network use a graph convolutional network (GCN). The graph convolutional network can effectively predict node signals with graph dependencies. Each layer of the graph convolutional network uses a graph filter, which embeds potential dependencies into a set of nodes that can aggregate input signals from other connected nodes. In this way, the output data of the graph convolutional network will show strong correlations between interconnected nodes.
[0090] Design a graph filter to match the actual spatial correlations between N photovoltaic power stations. Let L be the total number of hidden layers in the generator, be the input data matrix with features for each layer, where
[0091] the input of the first layer is a random noise matrix, that is, X (1) = Z and K1 = K. To gradually increase the time dimension, set For each layer use the matrix As a trainable feature filter parameter, it can be expressed as:
[0092]
[0093] In the formula, σ(·) is a non-linear activation function; in this embodiment, ReLU is selected as the activation function for most layers, and the sigmoid function is used for the last layer to generate output values in the range of [-1, 1] as the photovoltaic output data. The final output layer can present the correlation between the features of each node.
[0094] Each layer of the discriminator is similar to each layer of the generator. The first layer uses the sample data or real data generated by the generator as the input. The difference of the graph filter A is that the number of features decreases as increases, because this is a binary classification task, and the last layer outputs a binary decision. The activation function of the remaining layers all uses Leaky ReLU, and the sigmoid function is selected for the last layer to generate values in the range of (0, 1) to complete the binary classification.
[0095] Both the generator and the discriminator include graph filters in the spatial dimension and feature filters in the time dimension. In the spatial dimension, graph filters are used to capture the spatial correlation relationships between photovoltaic sites. Through the feature aggregation of adjacent nodes by graph convolution, the correlation coefficient after exponential transformation is used as the weight of the graph filter, as shown in Equation (7). This exponential transformation can improve the significance of the filter weights between relevant positions.
[0096]
[0097] In the time dimension, one-dimensional convolution is used to process the time series features to reduce the parameter scale and maintain the time-order features of the photovoltaic output. The matrix multiplication term in Equation (6) is regarded as the output of the feature filter. A one-dimensional convolution filter with a length of and trainable weight coefficients is used. The elements of the j-th column are as shown in the formula:
[0098]
[0099] In the formula, is the input data matrix with K l features for each layer, where l ∈ {1,..., L}, and L is the total number of hidden layers in the generator; σ(·) is a non-linear activation function; A is a graph filter, is a weight matrix; C i,j is the correlation coefficient between node i and node j; m is the position offset of the sliding window in the one-dimensional convolution; Ml is half of the width of the convolutional window; is the position feature matrix for the m offset; is the feature of the j-th column of the input matrix at the m offset position.
[0100] The above settings can match the temporal correlation of the photovoltaic output scenarios and reduce the complexity of the training process simultaneously.
[0101] After the above data preprocessing, random sampling noise is used as the input of the graph convolutional generative adversarial network generator, and the historical output data of each photovoltaic site is used as the input of the discriminator. Then, the graph filter in the spatial dimension and the one-dimensional filter in the temporal dimension are calculated respectively using equations (7) and (8). Finally, the graph convolutional generative adversarial network is trained using the training data until the Nash equilibrium is reached. After the training of the generative adversarial network is completed, a photovoltaic output scenario similar to the real data generated by the generator is obtained.
[0102] II. Reduction of photovoltaic output scenarios based on spatio-temporal clustering.
[0103] (III). Selecting typical photovoltaic output sites using ST-DBSCAN spatial clustering;
[0104] In this embodiment, the ST-DBSCAN clustering algorithm is used to cluster the power data of the photovoltaic power station, distinguish the data with concentrated density and the scattered data, and select typical photovoltaic sites. The steps are as follows:
[0105] Step 1. Set the input neighborhood radius Eps, the minimum number minpts, and the temporal neighborhood Δt;
[0106] Step 2. Starting from an arbitrary unlabeled point, select a point p as the core point for scanning. Scan the neighborhood within the temporal neighborhood Δt of the core point p, and count the number of neighborhood points that satisfy the distance threshold in the spherical region Eps(s) with the core point p as the center and the distance neighborhood as the radius within the temporal neighborhood Δt;
[0107] Step 3. If the number of neighborhood points of the core point p is less than the set threshold MinPts, mark this point as a noise point;
[0108] Step 4. If the number of neighborhood points is greater than the set threshold MinPts, mark this point as the core point p, generate a cluster number as cluster C1, and add all the points of the core point p in the neighborhood Eps(s) to the cluster C1 to be scanned;
[0109] Step 5. For the unlabeled points within the cluster C1, repeat steps S42 - S44 until all the points within the cluster C1 are marked;
[0110] Step 6: Continue to select the marked points from the dataset, repeat Steps 2 to 5, and cluster them into another class until all points are marked.
[0111] (4) Data reconstruction;
[0112] As Figure 3 shown, splice the outputs of the typical PV sites obtained by spatial clustering according to the corresponding dates into a one-dimensional sequence.
[0113] Splice the output of the typical site i obtained by spatial clustering into a matrix in the time dimension. Then, the 24-hour output data of all sites on the d-th day are spliced into a vector as:
[0114] V d = [P1(1,d), P1(2,d),..., P1(t,d),..., P i (1,D), P i (2,D),..., P i (t,D)] T (9);
[0115]
[0116] In the formula, P i (t,D) is the output of the i-th typical site at the t-th hour on the D-th day; t = 1, 2,..., 24, d = 1, 2,..., D, D is the total number of days, D = 365; i = 1, 2,..., M, where M is the number of typical sites; is a 24×M matrix.
[0117] (5) Select typical output dates using time clustering;
[0118] Use the K-medoids clustering method to select typical output dates. The specific steps are as follows:
[0119] Step 1: Randomly select r sequences from all sequences as the initial clustering centers, denoted by .
[0120] Step 2: According to the principle of the nearest distance to the clustering center, assign the remaining objects to each class, and calculate the distance between the two sequences with the nearest distance;
[0121]
[0122] In the formula, d(u i , u j ) is the distance between two sequences. For each sequence u i , select the clustering center u j with the smallest distance to it.to calculate the distance; T is the number of time points in the sequence; for scenario u i at the t-th time point; the distance d(u i , u j ) is the sum of the absolute values of the differences between the two sequences at each time point.
[0123] Step 3: The goal is to minimize the sum of the weighted distances from all scenarios to the cluster centers. According to the principle of minimizing the total loss of formula (11), find new cluster centers to replace the original cluster centers.
[0124]
[0125] In the formula, S is the set of all sequences; J is the set of cluster centers. The purpose is to find an optimal subset J of S to replace S, so that J contains as much information of S as possible; p i is the occurrence probability of sequence u i .
[0126] Step 4: Determine whether it converges. If it does not converge, repeat Step 2. If it has converged, then the obtained cluster centers are the typical days of PV output;
[0127] Step 5: Calculate the occurrence probability p i of each sequence after clustering. That is, p i is the ratio of the number of scenarios in the i-th class to the total number of scenarios.
[0128] The present invention believes that the complete output situation of each site corresponding to the typical date is the typical scenario of PV output. After obtaining the typical output date, in the returned generated output data, select the complete output slice corresponding to the output date, which is the typical scenario of PV output.
[0129] Embodiment 2
[0130] As Figure 5 shown, the present invention also provides a distributed PV random output scenario generation system, including:
[0131] Data acquisition module: used to acquire the historical output data of each PV site and perform normalization and standardization processing on the historical output data.
[0132] Graph convolutional generative adversarial network processing module: Use the noise as the input of the generator, and use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the graph convolutional generative adversarial network processing module to train the graph convolutional generative adversarial network.
[0133] Input the noise into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios, thereby completing the data augmentation of the output scenarios.
[0134] Spatial clustering processing module: Use the ST-DBSCAN clustering algorithm to cluster the power data of the photovoltaic power station, distinguish the data with concentrated density from the scattered data, select typical photovoltaic sites, and splice the outputs of each typical photovoltaic site according to the time dimension.
[0135] Temporal clustering processing module: Use the K-medoids clustering method to select typical output dates, and select the generated output scenarios from the completely generated photovoltaic output scenarios according to the dates to obtain typical photovoltaic output scenarios.
[0136] The distributed photovoltaic random output scenario generation system of the present invention can be installed in a computer device. The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a distributed photovoltaic random output scenario generation program. Among them, the memory includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The processor is the control core of the electronic device, connecting various components of the entire computer device through various interfaces and lines, and by running or executing the programs or modules stored in the memory, and calling the data stored in the memory, to execute various functions of the computer device and process data.
[0137] The module in the present invention refers to a series of computer program segments that can be executed by the processor of the computer device and can complete fixed functions, and are stored in the memory of the computer device.
[0138] The present invention uses a graph convolutional generative adversarial network to track and simulate the photovoltaic output characteristics, focuses on the internal laws of distributed photovoltaic output, explores the spatio-temporal correlation of the outputs of multiple distributed photovoltaic units, and generates the combined output scenarios of multiple units. On the basis of considering the randomness and correlation of the outputs of distributed photovoltaic units, artificial data with the same dimension and similar probability distribution as the original photovoltaic output data is generated, making the distributed photovoltaic output scenarios more comprehensive and diverse, and at the same time improving the accuracy of the distributed photovoltaic output scenarios. Using the scenario reduction method that clusters from two dimensions of space and time, the spatial correlation and temporal correlation between the outputs of each photovoltaic site are retained, typical photovoltaic sites and typical photovoltaic output days are extracted, and typical output scenarios with multiple photovoltaic sites are obtained.
Claims
1. A method for generating a distributed photovoltaic random output scenario, characterized in that, It includes the following steps: S1. Collect historical output data of each photovoltaic site, and perform normalization and standardization processing on the historical output data; S2. Use noise as the input of the generator in the graph convolutional generative adversarial network, use the generator to generate output data, use the output data generated by the generator and the processed historical output data as the input of the discriminator in the graph convolutional generative adversarial network, and train the graph convolutional generative adversarial network until reaching Nash equilibrium; S3. Input the noise into the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios, and complete the data augmentation of the output scenarios; S4. Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish the data with concentrated density and the scattered data, select typical photovoltaic sites, and splice the outputs of each typical photovoltaic site according to the time dimension; S5. Use the K-medoids clustering method to select typical output dates, and select the generated output scenarios from the completely generated photovoltaic output scenarios according to the dates to obtain typical photovoltaic output scenarios.
2. The method for generating a distributed photovoltaic random output scenario according to claim 1, wherein In S1, the historical output data of the PV sites are divided into the rainy season and the dry season, and the data format is set as a three-dimensional matrix. The PV power information of all sites for each day is recorded in the form of a two-dimensional matrix .
3. A method for generating a distributed photovoltaic random output scenario according to claim 1, characterized in that In the generator and discriminator of the graph convolutional generative adversarial network, the layers all use graph convolutional networks. The generator and discriminator both include graph filters in the spatial dimension and feature filters in the time dimension. In the spatial dimension, graph filters are used to capture the spatial correlation relationships between photovoltaic sites; in the time dimension, one-dimensional convolutional filters are used to process the time series features to complete feature extraction.
4. A method for generating a distributed photovoltaic random output scenario according to claim 3, characterized in that It includes the following steps: Use the input data matrix as the parameter of the feature filter: (6); The last layer uses the sigmoid function to generate output values in the range of (0,1) as the photovoltaic output data; Through the feature aggregation of adjacent nodes of graph convolution, use the exponentially transformed correlation coefficient as the weight of the graph filter, and the calculation expression is: (7); Regarding the matrix multiplication term in Equation (6) as the output of the feature filter, and using a one-dimensional convolutional filter with a length of and trainable weight coefficients, the elements in the th column are: (8); In the formula, is the input data matrix with features per layer, where , is the total number of hidden layers in the generator; is the non-linear activation function; is the graphical filter, , is the matrix; is the weight matrix, is the matrix; is the correlation coefficient between node and node ; is the position offset of the sliding window in the one-dimensional convolution; is half the width of the convolution window; is the weight of the convolution kernel; is the feature of the -th column of the input matrix at the offset position.
5. A method for generating a distributed photovoltaic random power output scenario according to claim 3, characterized in that, The loss function of the generator is: (3); The loss function of the discriminator is: (4); Wherein, is a noise variable; is a simple distribution; is the mapping of the noise variable to a generated data space; is the probability distribution of the real samples; is the generated sample; is the generator; is the discriminator; is the expectation.
6. A method for generating a distributed photovoltaic random power output scenario according to claim 1, characterized in that, In S4, the steps for selecting typical photovoltaic sites include the following: S41. Set the input domain radius , the minimum number and the time neighborhood ; S42. Select a certain point as the core point of scanning starting from any unlabeled point , the core point time neighborhood scan the neighborhood within the time neighborhood, scan the time neighborhood The number of neighborhood points that satisfy the distance threshold within the spherical region centered at the core point with the distance neighborhood as the radius ; S43. The number of neighborhood points of the core point is less than the set threshold , then mark this point as a noise point; S44. The number of neighborhood points is greater than the set threshold , then mark this point as a core point , generate a clustering number as cluster , and add the points in the neighborhood to the cluster to be scanned ; S45. For the cluster Repeat steps S42 to S44 for the unmarked points within the cluster until all the points within the cluster are marked; S46. Continue to select marked points from the data set, repeat steps S42 - S45, and cluster them into another class until all points are marked.
7. A method for generating a distributed photovoltaic random output scenario according to claim 6, characterized in that, The specific method of splicing the outputs of each typical photovoltaic site according to the time dimension is: splice the typical photovoltaic sites according to the corresponding dates to form a one-dimensional sequence: (9); ; In the formula, is the output of the th typical site at the th day and the th hour; , , is the total number of days, ; , where is the number of typical sites; is a matrix of .
8. A method for generating a distributed photovoltaic random output scenario according to claim 1, wherein S5 includes the following steps: S51. Randomly select sequences as the initial clustering centers; S52. According to the principle of the closest distance to the cluster center, assign the remaining objects to each class, and calculate the distance between the two closest sequences; S53. Minimize the weighted distance sum of all scenarios to the cluster center; S54. Determine whether it converges. If it does not converge, repeat S52. If it has converged, then the cluster centers obtained by clustering are the typical days of PV output; S55. Calculate the probabilities of each sequence appearing after clustering.
9. A method for generating a distributed photovoltaic random power output scenario according to claim 8, characterized in that, The calculation expression for calculating the distance between the two closest sequences is: (10); wherein, is the distance between two sequences, and for each sequence , the clustering center with the minimum distance thereto is selected to calculate the distance; is the number of time points of the sequence; is the scene at the -th time point; the distance is the sum of the absolute values of the differences between the two sequences at each time point.
10. A distributed photovoltaic random output scenario generation system, which applies the distributed photovoltaic random output scenario generation method described in any one of claims 1-9, is characterized in that, It includes: Data acquisition module: used to acquire historical output data of each photovoltaic site, and perform normalization and standardization processing on the historical output data; Graph convolutional generative adversarial network processing module: use noise as the input of the generator, use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the graph convolutional generative adversarial network processing module, and train the graph convolutional generative adversarial network; Input the noise into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios and complete the data augmentation of the output scenarios; Spatial clustering processing module: Use the ST-DBSCAN clustering algorithm to cluster the power data of the photovoltaic power station, distinguish the data with concentrated density from the scattered data, select typical photovoltaic sites, and splice the outputs of each typical photovoltaic site according to the time dimension; Temporal clustering processing module: Use the K-medoids clustering method to select typical output dates, and select the generated output scenarios from the completely generated photovoltaic output scenarios according to the dates to obtain typical photovoltaic output scenarios.
Citation Information
Patent Citations
Photovoltaic power station typical scene generation method based on multi-scene model
CN112541546A
Wind-solar active power output scene generation method and device, electronic equipment and storage medium
CN114066236A